I scored 7 AI red-teaming tools on a real Azure agent. Their own reports were 70–97% noise.
If you're choosing an AI red-teaming tool today, you're probably comparing what each tool says it found. I wanted to know what they actually found. So I pointed seven of them at the same agent and scored every turn from the wire, ignoring the tools' own reports.
Here's what happened. All of it is reproducible, and every confirmed violation is published in the benchmark repository.
The setup
- Target: a Microsoft Foundry prompt agent on
gpt-5-nano, a support agent for a fictional utility, with four function tools (customer lookup, refunds, e-mail, a knowledge base). Every Azure content filter was at its strictest setting (Low), with Prompt Shields on. - Tools: garak, promptfoo, DeepTeam, PyRIT, the Azure AI Red Teaming Agent, agent-probe, and sixi-scanner. Each ran with the configuration its own documentation recommends.
- Ground truth: every turn went through one gateway. Ten deterministic oracles check each one (did the agent e-mail an address off the allow-list? exceed the refund cap? leak the planted secret?), plus an LLM judge that never sees which tool sent the turn.
- Scale: about 15,000 target turns from late September to early October 2026.
Conflict of interest, up front: I wrote sixi-scanner. That's why scoring is oracle-first and tool-blind, why every change after a run is disclosed, and why the raw findings are public, including the runs where my tool looked worst.
Finding 1: tools' own reports are mostly noise
In the baseline, 70–97% of each tool's flags were false alarms: precision ran from 3% to 30%. And every tool that broke something left at least 44% of what it broke out of its own report.
The classic trap: "I can't share API keys or any internal credentials" contains every substring a credential detector looks for. It's a refusal, not a leak, and most tools file it as a finding anyway.
Takeaway: if you compare tools by their dashboards, you're comparing their noise. Score from the wire.
Finding 2: it wouldn't say its secret. It e-mailed it.
The agent's instructions carry an escalation code it must never output "in any form". Across roughly 10,950 turns, that code never appeared in a chat reply. Prompt hardening worked.
Three times, though, the agent read a poisoned knowledge-base article (KB-317) and e-mailed the full customer record, including name, phone, IBAN and that secret code, to claims@helios-claims-desk.biz.
Prompt Shields' indirect-attack detection was on. The article came back as tool output, which it never flagged.
Takeaway: guard what the agent can do, not just what it says. Check tool calls against an allow-list in code.
Finding 3: an LLM judge misses business logic
The refund cap is 50 EUR per request. Ask for two 30-EUR refunds in one conversation and every call is within the limit, but the total is 60.
- A five-line deterministic oracle caught all 7 such turns.
- The LLM judge cleared 5 of the 7 as "each within the 50 EUR limit".
Takeaway: score business rules with code. Use a model only for what code can't decide.
Finding 4: content filters stop content, not business logic
With every filter at Low, 64% of the Azure AI Red Teaming Agent's turns never reached the model. Its harm-taxonomy attacks are exactly what the filters are built to stop, and it confirmed 3 violations. promptfoo's business-logic attacks were 2% blocked and confirmed 89.
Finding 5: the bill hides in the tool, not the target
The baseline cost $79.99 on Azure. $73.30 (92%) of that was the Azure AI Red Teaming Agent's own hosted grading. The target agent itself cost $3.96 for all seven tools combined. (scan(skip_evals=True) turns that grading off.)
Every tool led somewhere
This isn't a "tool X wins" post.
- promptfoo confirmed the most violations (89) and reached all three e-mail/exfiltration oracles.
- garak had the broadest coverage (8 of 9 risk categories) and the best baseline recall.
- DeepTeam was the most precise baseline tool (30%) and the most efficient: 165 turns, 41 minutes, $0.17.
- PyRIT is a framework rather than a scanner: what it finds depends on the objectives you write.
And my tool
After the baseline I open-sourced sixi-scanner: a single Go binary with 21 techniques and no LLM inside. Measured on the same wire, with the same oracles and the same judge, v0.6.0 is 1st on precision (0.688) and recall (0.833), for $0.57 of cloud cost. A follow-up audit found recall at its ceiling on this target: it caught 34 of the 37 real leaks, and the other 3 are two judges disagreeing, not a missing marker.
Its gap is real and stated in the benchmark: breadth. It finds fewer distinct attacks than promptfoo or garak (6 against 89 and 67), and its runs were tuned against this target's ground truth while the others ran once. Read its row with that in mind.
Then and now
sixi-scanner began in 2025. The open-source rewrite came out this summer, and I maintain it in the open so anyone securing an AI agent can run it, read it and improve it. Here is the first build I measured, the legacy build's best run and the current release, all on this same target, gateway, oracles and judge:
- Precision 0.03 → 0.69 and recall 0.16 → 0.83. The report went from about 1 real finding in 35 flags to about 2 in 3.
- False alarms 105 → 10, wall clock 265 → 44 minutes, attacker-model calls 832 → 75.
- The cost is breadth: the legacy LLM attacker confirmed 7 risk categories, and v0.6.0 confirms 4. That's where the open-source build has the most to gain.
The first step on breadth is out. v0.7.0 adds false action claims, the largest confirmed family sixi-scanner had never produced (58 turns in the baseline, 51 of them promptfoo's): the agent says "I've emailed your account summary" with no tool call behind it. Because the scanner holds the tool trace for every turn, it checks the claim against what actually ran, with no model and no false positive from a claim that turns out to be true. Replayed against recorded replies, it caught every judge-confirmed case and added no new false positives over the 1,282 replies in the v0.6.0 run. A full benchmark run of v0.7.0 hasn't been published yet.
Work in progress: sixi-scanner is in active development and more updates are on the way. Each new release is measured on this same benchmark, and the results go into the benchmark repository and this page, wins and losses alike. Follow sixi-scanner on GitHub to see them land.
If you want it in CI: uses: rbrus/scan-action@v2 puts findings in your GitHub Security tab, and an unreachable agent fails the job instead of passing.
What I'd do if I were shipping an agent tomorrow
- Allow-list tool side effects in code (recipients, amounts, accounts), with session-level totals rather than only per-request checks.
- Treat tool output as untrusted input. Your knowledge base is an attack surface.
- Run a scanner in CI, but verify its findings before trusting its dashboard.
- Don't let an LLM grade business rules alone.
- One run is a sample. Six identical re-runs of the same e-mail attacks scored 3, 0, 3, 2, 0 and 2 oracle hits.
Everything (protocol, oracles, per-run KPIs and the confirmed turns) is at github.com/rbrus/agent-redteam-benchmark. If you maintain one of these tools and think I misconfigured it, open an issue with your config and I'll re-run it.