← AI Switchboard
AI Switchboardby Waggle
RESEARCH · September 28, 2026
Sep 28

A two-agent auditor finds real bugs in 12 open-source agent frameworks, and beats Codex at it

AgentXploit splits pre-deployment auditing into a repository analyst and a runtime exploiter, reporting 59.3% end-to-end success against 38.4% for Codex on a new benchmark of 72 reproducible vulnerabilities. Give Codex a matched token budget and it reaches 46.3%, which narrows the gap without closing it.

The setting is authorised white-box pre-deployment auditing: the auditor has the target repository and a controlled runtime. The detail that makes the numbers mean something is that any successful attack must still act through the task-defined attacker interface and be confirmed by an external verifier. That external confirmation is what separates a real exploit from a model asserting it succeeded, and it is the protocol to name whenever this 59.3% is quoted.

The architecture reflects a claim about where the difficulty lives. An Analyzer Agent traces attacker-controlled inputs through to sensitive operations and records candidate attack paths supported by the code; an Exploiter Agent turns those paths into concrete attacks and revises them from runtime feedback. The conclusion is that repository-level discovery and runtime exploitation are distinct challenges — which is why the two comparisons behave differently. Where the attack path must be found, the system beats Codex 59.3% to 38.4%. On a second benchmark where injection points are already provided and only exploitation remains, the Exploiter Agent alone reaches 79.2% against 52.7% for the prior tool.

The token-budget control is the honest part of the paper and belongs beside the headline. Codex rises from 38.4% to 46.3% when given a matched budget, so a meaningful slice of the advantage is compute rather than architecture. What survives the control is about 13 percentage points. The vulnerability classes are ordinary software bugs — path traversal, command injection — reachable through the agent's data and tool surface, which is the point: the failure is as often in the software wrapped around the model as in the model.

Limits worth carrying. The benchmark is introduced by the same authors who report state-of-the-art on it, so the benchmark and the result have not yet been separated and nobody outside the group has run it. The targets are open-source agent frameworks, not hosted commercial products, so the numbers do not transfer to a claim about any deployed assistant. And three runs is a small sample for a twenty-point gap.

  • Confirmed AgentXploit reaches 59.3% end-to-end success against 38.4% for Codex across three runs; under a token-budget-matched comparison Codex reaches 46.3%. arXiv:2609.31318
  • Confirmed The benchmark contains 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. arXiv:2609.31318
  • Confirmed With injection points provided, the Exploiter Agent alone reaches 79.2% attack success against 52.7% for the prior tool. arXiv:2609.31318

Science & researchSafety, security & governance

Today in the September 28, 2026 edition · front page