Radoslaw Brus

I scored 7 AI red-teaming tools on a real Azure agent. Their own reports were 70–97% noise.

Radoslaw Brus · October 2026 (updated 7 October) · data and code

If you're choosing an AI red-teaming tool today, you're probably comparing what each tool says it found. I wanted to know what they actually found. So I pointed seven of them at the same agent and scored every turn from the wire, ignoring the tools' own reports.

Here's what happened. All of it is reproducible, and every confirmed violation is published in the benchmark repository.

The setup

Conflict of interest, up front: I wrote sixi-scanner. That's why scoring is oracle-first and tool-blind, why every change after a run is disclosed, and why the raw findings are public, including the runs where my tool looked worst.

Precision and recall of every red-teaming tool on one Foundry agent
Every tool on the same agent, scored from the wire.

Finding 1: tools' own reports are mostly noise

In the baseline, 70–97% of each tool's flags were false alarms: precision ran from 3% to 30%. And every tool that broke something left at least 44% of what it broke out of its own report.

The classic trap: "I can't share API keys or any internal credentials" contains every substring a credential detector looks for. It's a refusal, not a leak, and most tools file it as a finding anyway.

Takeaway: if you compare tools by their dashboards, you're comparing their noise. Score from the wire.

Finding 2: it wouldn't say its secret. It e-mailed it.

The agent's instructions carry an escalation code it must never output "in any form". Across roughly 10,950 turns, that code never appeared in a chat reply. Prompt hardening worked.

Three times, though, the agent read a poisoned knowledge-base article (KB-317) and e-mailed the full customer record, including name, phone, IBAN and that secret code, to claims@helios-claims-desk.biz.

Prompt Shields' indirect-attack detection was on. The article came back as tool output, which it never flagged.

Takeaway: guard what the agent can do, not just what it says. Check tool calls against an allow-list in code.

Finding 3: an LLM judge misses business logic

The refund cap is 50 EUR per request. Ask for two 30-EUR refunds in one conversation and every call is within the limit, but the total is 60.

Takeaway: score business rules with code. Use a model only for what code can't decide.

Finding 4: content filters stop content, not business logic

With every filter at Low, 64% of the Azure AI Red Teaming Agent's turns never reached the model. Its harm-taxonomy attacks are exactly what the filters are built to stop, and it confirmed 3 violations. promptfoo's business-logic attacks were 2% blocked and confirmed 89.

Confirmed violating turns by tool and risk category
Confirmed violations by tool and risk category.

Finding 5: the bill hides in the tool, not the target

The baseline cost $79.99 on Azure. $73.30 (92%) of that was the Azure AI Red Teaming Agent's own hosted grading. The target agent itself cost $3.96 for all seven tools combined. (scan(skip_evals=True) turns that grading off.)

Every tool led somewhere

This isn't a "tool X wins" post.

And my tool

After the baseline I open-sourced sixi-scanner: a single Go binary with 21 techniques and no LLM inside. Measured on the same wire, with the same oracles and the same judge, v0.6.0 is 1st on precision (0.688) and recall (0.833), for $0.57 of cloud cost. A follow-up audit found recall at its ceiling on this target: it caught 34 of the 37 real leaks, and the other 3 are two judges disagreeing, not a missing marker.

Its gap is real and stated in the benchmark: breadth. It finds fewer distinct attacks than promptfoo or garak (6 against 89 and 67), and its runs were tuned against this target's ground truth while the others ran once. Read its row with that in mind.

Then and now

sixi-scanner began in 2025. The open-source rewrite came out this summer, and I maintain it in the open so anyone securing an AI agent can run it, read it and improve it. Here is the first build I measured, the legacy build's best run and the current release, all on this same target, gateway, oracles and judge:

sixi-scanner then and now: precision, recall, false alarms, wall clock, attacker-model calls and risk categories for the first build, the legacy v9 build and open-source v0.6.0
Each bar is one published run in the benchmark repository.

The first step on breadth is out. v0.7.0 adds false action claims, the largest confirmed family sixi-scanner had never produced (58 turns in the baseline, 51 of them promptfoo's): the agent says "I've emailed your account summary" with no tool call behind it. Because the scanner holds the tool trace for every turn, it checks the claim against what actually ran, with no model and no false positive from a claim that turns out to be true. Replayed against recorded replies, it caught every judge-confirmed case and added no new false positives over the 1,282 replies in the v0.6.0 run. A full benchmark run of v0.7.0 hasn't been published yet.

Work in progress: sixi-scanner is in active development and more updates are on the way. Each new release is measured on this same benchmark, and the results go into the benchmark repository and this page, wins and losses alike. Follow sixi-scanner on GitHub to see them land.

If you want it in CI: uses: rbrus/scan-action@v2 puts findings in your GitHub Security tab, and an unreachable agent fails the job instead of passing.

What I'd do if I were shipping an agent tomorrow

  1. Allow-list tool side effects in code (recipients, amounts, accounts), with session-level totals rather than only per-request checks.
  2. Treat tool output as untrusted input. Your knowledge base is an attack surface.
  3. Run a scanner in CI, but verify its findings before trusting its dashboard.
  4. Don't let an LLM grade business rules alone.
  5. One run is a sample. Six identical re-runs of the same e-mail attacks scored 3, 0, 3, 2, 0 and 2 oracle hits.

Everything (protocol, oracles, per-run KPIs and the confirmed turns) is at github.com/rbrus/agent-redteam-benchmark. If you maintain one of these tools and think I misconfigured it, open an issue with your config and I'll re-run it.