The OWASP LLM Top 10, and guardrails
Enough teams shipped the same ten mistakes that someone numbered them. Most are ordinary application security with a more persuasive input.
Part of the Trust and evals track on lAItest.
Enough teams shipped the same ten mistakes that somebody wrote them down and gave them numbers.
It is the cheapest security review you will ever run.
The OWASP Top 10 for LLM applications, 2025
LLM01 Prompt Injection. LLM02 Sensitive Information Disclosure. LLM03 Supply Chain. LLM04 Data and Model Poisoning. LLM05 Improper Output Handling. LLM06 Excessive Agency. LLM07 System Prompt Leakage. LLM08 Vector and Embedding Weaknesses. LLM09 Misinformation. LLM10 Unbounded Consumption.
Read the list for its shape
Only a few of the ten are about the model being wrong. The rest are ordinary application security problems in new clothes: what you feed it, where its output goes, what it is permitted to do, what it leaks, and what it costs when a loop does not stop. LLM05 is worth spelling out — model output arriving somewhere that executes it, a shell, a query, a rendered page, is the same class of bug as unsanitised user input, with a far more persuasive source.
A common misconception
Commonly believed: Guardrails are a feature you switch on. Choose a guardrail product, enable it, and the application is safe.
Actually: A guardrail is anything that constrains the system, and the ones that hold are mostly not model settings: what the model is permitted to call, what a human must approve, what the output is allowed to touch, a cap on steps and spend so a loop cannot run all night. Model-side checks help. They are not sufficient, because a check the model performs on itself can be fooled by the same input that fooled the model.
An agent loops on a failing tool for six hours and burns the month’s budget. Which item on the list is that?
Answer: LLM10 Unbounded Consumption. Unbounded consumption is the cost and availability item: no ceiling on steps, tokens or spend. Excessive agency is nearby and often present too, but that one is about a model being allowed to take actions the task never needed. Both are fixed by a limit set outside the model, because a model cannot reliably notice that it is stuck.
In one sentence
Most LLM security failures are ordinary security failures with a more persuasive input, and most guardrails that work are limits set outside the model.