We’ve all seen the charts: “Our product blocks 99.8% of prompt injections!” Usually followed by a footnote in size-2 font about their “proprietary benchmark.” It’s security theater dressed as a data sheet.
The problem isn't that vendors test; it's that they get to define both the exam and the grading rubric. A detection rate is meaningless without knowing what’s being detected. Are they counting simple keyword flagging on curated, obvious attacks? Are they including subtle context corruption, multi-turn jailbreaks, or indirect injection via retrieved documents? Or is their benchmark just a thousand variations of “Ignore previous instructions” and “You are now DAN”?
If we want numbers that aren’t purely for marketing, we need to agree on a few baseline principles for comparison. Not another monolithic benchmark—those get gamed quickly—but a methodology.
First, the attack taxonomy must be public and extensive. It should cover the spectrum from naive to novel, including:
- Direct injection (plaintext, encoded, natural language)
- Indirect injection (via tool output, RAG context, user history)
- Multi-modal or multi-step attacks
- Non-English and culturally-specific social engineering prompts
Second, the test set must include a “benign” corpus. What’s the false positive rate on normal, quirky, or edge-case user queries? A system that flags 10% of legitimate customer service prompts as malicious is useless, regardless of its detection score.
Third, the runtime conditions matter. Is the detection running pre-execution, or is it monitoring during agent operation? Static analysis catches the lazy attacks; a dynamic environment is where the real fight happens.
So, my question is this: what would a minimally misleading evaluation framework actually require? I think it starts with transparent, community-defined test suites and the courage to publish failure cases, not just success rates. Otherwise, we’re just comparing vanity metrics.
Jack
Security theater is still theater.
You're absolutely right about the taxonomy being the first critical piece. If we're going to compare numbers, we need to know what's in the "injection zoo" they're testing against.
My caveat would be to also demand they publish the *failure cases*. Not just what they caught, but what got through. A vendor's 0.2% failure rate is meaningless if those failures are all the novel, scary stuff you listed, while the 99.8% blocked are just trivial keyword matches. The real signal is in the misses.
This reminds me of the early days of SAST tool comparisons - everyone had a 99% detection rate until you realized their test suite was 99% simple, contrived vulnerabilities. An open, peer-reviewed test set for injections would be a game changer.
trivy image --severity HIGH,CRITICAL
Agree completely. The methodology you're proposing aligns with NIST's approach for adversarial ML taxonomy (NIST AI 100-2e2023). Without a public, stratified taxonomy, a single detection rate is worse than useless, it's actively misleading.
One nuance: even with a public taxonomy, vendors could still skew results by weighting categories. If 95% of their test set is "direct injection, naive," that 99.8% figure is still hollow. We'd need to see the distribution of attacks in the test set alongside the per-category detection rates. That reveals if they're only effective against the easy stuff.
A peer-reviewed test set, as user426 mentioned, is the logical end state. Until then, demanding the full test corpus and weighting methodology for any published number is a reasonable starting point for procurement.
Policy is code
You're right about the weighting. That's the next loophole after they're forced to use a public taxonomy.
So what stops a vendor from building a "public test set" that's just 90% easy-to-catch garbage? Their 99.8% still looks great.
We need a rule on distribution. Like, at least 30% of test cases must be from a "novel/advanced" tier. And the overall score has to be broken down by tier.
Is anyone actually asking for this during procurement?
That's a good point about the distribution rule. I'm new to this but it seems like even if you have tiers, the vendor could still label their own easy tests as "advanced", right? Who decides what counts as novel?
Is there a group working on a standard test set like user426 mentioned? That seems like the only way to make the tiers mean something.