Fundamentals

OWASP Top 10 for LLM Apps: Turning It Into Controls

Map each OWASP LLM Top 10 risk to an owner, a preventive control, a monitoring signal and a test, so the list becomes evidence instead of a slide.

4 min read
OWASP Top 10 for LLM Apps: Turning It Into Controls

A risk list tells you what to worry about. It does not tell you what you have done about it. When a security review asks for evidence that your AI feature is protected, "we read the OWASP Top 10" is not an answer. A risk-to-control matrix is.

The OWASP Top 10 for LLM applications (2025 edition) contains ten entries. This article turns each into something a team can assign, build and test.

The matrix

OWASP risk Preventive control Detection or monitoring signal Example test
LLM01 Prompt Injection Separate trusted instructions from untrusted content; limit what a manipulated model can reach Input and retrieved-content scan scores; blocked-event rate by category Versioned attack corpus, direct and indirect (testing checklist)
LLM02 Sensitive Information Disclosure Data minimization; authorization before retrieval; redaction Sensitive-data flags at output; egress alerts Prompt that asks for another user's data; canary secrets in context
LLM03 Supply Chain Approved model, dataset and plugin inventory; pinned versions; review of third-party servers Change alerts on models, packages and tool definitions Verify hashes; simulate a changed tool description (untrusted MCP servers)
LLM04 Data and Model Poisoning Control who can write to training and knowledge sources; provenance and versioning Anomalies in new content; retrieval pattern shifts Insert a marked document and confirm it is quarantined (RAG data poisoning)
LLM05 Improper Output Handling Validate and encode output for each destination; parameterized APIs Blocked-output events; sink errors Feed script tags, URLs and commands through the output path (output pipeline)
LLM06 Excessive Agency Least-privilege tools and credentials; approvals for high-impact actions Tool-call logs; approval and denial rates Attempt an out-of-scope action and confirm it fails (excessive agency)
LLM07 System Prompt Leakage Keep secrets and authorization out of prompts Leakage flags in outputs; repeated probing Extraction attempts against staging (system prompt leakage)
LLM08 Vector and Embedding Weaknesses Access control in retrieval; ingestion validation; tenant isolation Unusual retrieval patterns; cross-tenant checks Cross-tenant query returns nothing (vector database security)
LLM09 Misinformation Grounding in trusted sources; citations; human review where errors are costly User feedback; spot audits Golden question set with known answers
LLM10 Unbounded Consumption Rate limits, quotas, token and cost caps, timeouts Spend and token dashboards; per-user usage Load test and abuse simulation

Treat the table as a starting point. The right control depends on your architecture: a chatbot without tools has little exposure to LLM06, and a batch classifier with no retrieval may have none to LLM08.

How to use it in practice

  1. Scope the review to your architecture. List your inputs, data sources, model calls, output destinations and tools. Mark which of the ten risks each one touches.
  2. Assign one owner per control. A control with no owner decays. Owners can be different teams: platform for rate limits, data for ingestion, application for output handling.
  3. Attach a test and a piece of evidence. "Cross-tenant queries return no chunks" is a testable statement with a test result as evidence. "We take tenancy seriously" is not.
  4. Re-review on change. Add a model, a data source, a tool or a permission and the mapping may shift. Tie the review to those changes, not to a calendar alone.

The same approach scales into a launch review, covered in the production GenAI security checklist.

What no single control covers

Some risks are model-facing, some are classic application security, and some are supply-chain concerns around AI. One control can help several risks. Least privilege, for instance, limits prompt injection, excessive agency and data disclosure at once. But no single scanner covers the whole list. A runtime scanner is useful evidence for LLM01, LLM02, LLM05 and LLM07, and says little about supply chain, cost or misinformation. See what an LLM firewall is and is not.

Common mistakes

  • Claiming "OWASP compliant" without defining scope. The list is guidance, not a certification.
  • Using one vendor feature as evidence for all ten risks.
  • Skipping ordinary secure development practices because "the model handles it".

Sources and further reading

Frequently asked questions

What is the OWASP Top 10 for LLM applications?

It is a community-maintained list of the ten most critical risks for applications built on large language models, including prompt injection, sensitive information disclosure, improper output handling and excessive agency. The 2025 edition is the current reference used in this article.

How do I turn the OWASP list into something I can audit?

Build a matrix that maps each relevant risk to a preventive control, an owner, a monitoring signal and a test, then keep the test results as evidence. Review the matrix whenever models, data sources, tools or permissions change.

Does one security tool cover the whole OWASP LLM Top 10?

No. A runtime scanner helps with several risks such as prompt injection and data disclosure, but supply chain, unbounded consumption and misinformation need other controls, including inventory, quotas and grounding.

Is there an OWASP LLM certification?

No. Teams can map their controls to the list and demonstrate coverage, but there is no compliance badge attached to it, so avoid claiming "OWASP compliant" without stating scope.

Keep reading

Evaluation & Ops

AI Security Logging and Observability

What to log for AI security events and what to leave out: request IDs, categories, policy versions and tool calls, with raw prompts kept separate.

#logging #monitoring4 min read