Prompt-injection detectors that ace one benchmark can miss almost everything in your agent
A 2 October 2026 arXiv study of fifteen detectors finds benchmark rankings barely transfer to agent tool outputs: one detector catches 95.1% on BIPIA but 2.1% on AgentDojo.
What is this?
Prompt-injection detectors are small classifiers placed between an agent and the content it reads, such as web pages, emails and tool results. They flag text that tries to give the agent new instructions. Vendors usually justify them with a score on a public benchmark. A preprint submitted to arXiv on 2 October 2026 (arXiv:2610.03448, “Passing the Test You Trained On”, single author Zhuowen Liu) asks a practical question: does that score predict how the detector behaves inside an agent?
The answer is largely no. The study evaluates fifteen detectors plus two LLM judges and finds that rankings on a text-oriented benchmark (BIPIA, published December 2023) say almost nothing about performance on agent tool outputs (AgentDojo, June 2024, and tau-bench).
How it works
The authors replay the same ground-truth tool calls in a clean environment and in an injected one. Comparing the two outputs, a method they call differential replay, lets them label which tool outputs actually contain the injection. This avoids brittle substring matching, which breaks when tool results are re-serialized as YAML or JSON.
They then measure detection at a fixed 1% false-positive rate on three sets: AgentDojo v1.2 (339 clean and 1,152 injected outputs), tau-bench (1,533 clean and 584 injected outputs) and BIPIA (1,200 benign and 2,435 injected samples).
The headline gaps are large. One detector catches 95.1% of BIPIA injections but only 2.1% of AgentDojo injections. Another catches 72.2% of AgentDojo injections but only 15.2% on tau-bench. Rank correlation between BIPIA and AgentDojo results is essentially zero (Kendall’s tau 0.01). Between the two agent benchmarks it is weak (0.28). False-positive behaviour, by contrast, is stable across agent benchmarks (tau 0.67), and it varies enormously: from 0% to over 90% of benign tool outputs flagged, depending on the detector.
The best performer across agent benchmarks was trained on agent-style inputs and shared no data with the benchmarks it was tested on, which points to training-data distribution, not leaderboard rank, as the driver.
Why it matters
A detector tuned on short natural-language injections can look excellent while missing injections embedded in structured tool output. Conversely, a detector that flags a large share of ordinary JSON results will break agent workflows and be switched off. Both failure modes stay invisible if you only read a headline accuracy figure.
The authors note limits: the environments are simulated, the attacks are static templates from public benchmarks, and adaptive attackers would likely lower detection further. Only two detectors’ training data could be audited. Treat the numbers as a warning about methodology, not as a ranking of products.
Defenses
Evaluate detectors on your own agent’s real tool outputs, not on generic text. Report detection at a fixed low false-positive rate such as 1%, and measure false positives separately on benign outputs from your tools, since this is the number that transfers. Check whether a detector was trained on the benchmark it is quoted on. Prefer detectors trained on agent-style inputs, and re-test after every change to tool formats.
Do not rely on a detector as the only control. Pair it with architectural limits: least-privilege tool permissions, human confirmation for high-impact actions, separation of untrusted content from instructions, and logging of tool calls for review.
Status
| Item | Detail |
|---|---|
| Paper | arXiv:2610.03448, submitted 2 October 2026 |
| Scope | 15 detectors + 2 LLM judges; AgentDojo, tau-bench, BIPIA |
| Key result | 95.1% on BIPIA vs 2.1% on AgentDojo for the same detector (1% FPR) |
| Limitations | Simulated environments, static attacks, limited training-data audit |
| Fix status | Methodology guidance; no product patch involved |