A CVSS score is a property of a vulnerability. “Exploitable here” is a property of a path. Panop’s autonomous pentest exists to turn the second into evidence: a finding is confirmed only when the platform has demonstrated access, under guardrails, with the reproduction path attached.
That claim is only as good as the targets you test it on. Scanner labs you wrote yourself do not count. We are scoring Panop against three public, independently authored suites that were built to evaluate offensive agents, not to flatter a vendor:
- XBOW Validation Benchmarks — 104 Jeopardy-style web challenges.
- Argus Validation Benchmarks — 71 live pentest targets on modern stacks, plus a separate threat-model suite.
- ExploitBench — a 16-flag exploitation ladder on 41 Chromium V8 bugs.
This page is the protocol. The results tables below stay empty until the run completes. We will fill them in place, not in a separate announcement that can drift from the method.
Why three suites, not one?
Each suite asks a different question. Treating them as one leaderboard would hide that.
| Suite | What it asks | Win condition | Why it is in this evaluation |
|---|---|---|---|
| XBOW | Can the agent find a hidden flag in a web CTF target? | Flag | The 2024–25 reference set for web offensive tools. XBOW now marks it as saturated (~100% industry performance as of mid-2026). We run it as a baseline, not as the discriminating test. |
| Argus | Can the agent exploit modern, multi-step, sometimes defence-enabled applications? | Flag | Built because older sets over-index on PHP and under-cover GraphQL, race conditions, JWT, WAF bypass, sandbox boundaries and attack chains. This is the discriminating application suite. |
| ExploitBench | How far up a real exploitation ladder does the agent climb? | Capability bitmap, graded by deterministic oracles | Crash-only benches stop at reproduction. ExploitBench scores coverage → trigger → primitives → arbitrary code execution on production V8 with the heap sandbox on. This is the discriminating engine suite. |
A high XBOW score, in 2026, says the agent can still do the work that used to separate tools. It does not say it can pentest a GraphQL SaaS app, bypass a WAF, or turn a V8 N-day into a sandbox escape. That is why Argus and ExploitBench are here.
What XBOW measures
XBOW Validation Benchmarks are 104 isolated challenges, each with one objective: recover a flag from a Dockerised web target. Flags are injected at build time, not hardcoded. Difficulty is author-assigned as level 1 (easy), 2 (medium) or 3 (hard). The original brief was to mirror vulnerability classes seen in pentest and bug-bounty work, and to evaluate web-based offensive tools rather than trivia.
XBOW’s own README now warns that the set is outdated for ranking. Released in 2024, with build fixes in 2025, it has been trained into current models. As of mid-2026 the maintainers put industry-standard performance at approximately 100%, and they keep the repository for historical comparison.
We still run it, for two reasons. First, most published agent papers from 2024–25 used it, so a Panop number that cannot be placed next to those papers is not useful. Second, a miss on a saturated set would be a real signal. A perfect score is not.
Panop is scored on flag capture. A challenge counts as solved only when the injected flag is recovered. We will report solved / 104 and a breakdown by level.
What Argus measures
Argus Validation Benchmarks, published by Pensar, are the most complete open pentest set we could find that is not already spent. Two suites live in the same repository. Only the first is in scope for this product evaluation.
Pentesting benchmarks (71) — self-contained Docker applications and infrastructure targets. Each has a planted vulnerability that must be discovered and exploited to capture a flag in the form FLAG{…}. Coverage includes Node.js, Python, Java, Go, PHP and Ruby; injection, authz bypass, SSRF, XSS, prototype pollution, race conditions, deserialization, SSTI, WAF/IDS bypass, cloud and CI/CD chains; and, in the 2026 ARGUS-*-26 family, sandbox escapes, OT/ICS, firmware, identity, ML supply chain, mobile and edge.
Difficulty, as assigned by the authors:
| Difficulty | Count | Their description |
|---|---|---|
| 1 (Easy) | 2 | Single request, obvious vulnerability, standard payload |
| 2 (Medium) | 27 | Multiple requests, some enumeration, encoded payloads |
| 3 (Hard) | 42 | Multi-step, timing-sensitive, chaining or boundary reasoning |
Threat-model benchmarks (10) — source-code analysis against a ground-truth document, scored on six dimensions (structure, grounding, anti-patterns, discovery quality, attack-path depth, usefulness). That is a different product question than live exploit validation. It is out of scope for this run. We may score it later; we will not mix those numbers into the pentest tables.
Panop is scored on flag capture against the 71 live targets. A challenge counts as solved only when the suite’s own checker accepts the flag. We will report solved / 71, then splits by difficulty and by the authors’ vulnerability categories.
What ExploitBench measures
ExploitBench (Lee and Brumley, Carnegie Mellon; public results at exploitbench.ai) does not award a single flag. It asks how far an agent climbs, on a known V8 bug, against a production-like d8 with the V8 heap sandbox and shipped mitigations enabled.
The current instantiation, bench-v8, covers 41 V8 bugs reported no earlier than 2024. Each episode yields a 16-bit capability map, grouped into five tiers:
| Tier | Name | What a success means |
|---|---|---|
| T5 | Coverage | The patched function or line actually ran |
| T4 | Reproduction | Crash, sanitizer report, or differential behaviour versus the fixed build |
| T3 | Target primitives | V8-in-sandbox building blocks (addrof, fakeobj, caged read/write) |
| T2 | Generic primitives | Arbitrary read/write and infoleaks past the engine’s isolation |
| T1 | Full control | Control-flow hijack and arbitrary code execution |
Every rung is checked by a deterministic oracle compiled into the grader — coverage, differential execution, randomised challenge-response for primitives, a signal-handler proof for PC control. There is no LLM-as-judge. Crash-class benches (the authors name CyberGym, CyBench, SEC-bench Pro) stop at T4. The point of ExploitBench is the climb above that floor.
Public frontier models, in the authors’ own panel, routinely reach coverage and often crash. Arbitrary code execution on this set is rare. That is the reason to include it: if Panop’s “exploitable” only ever means “web flag captured”, we have not measured exploitation against a hardened engine.
Panop is scored with the suite’s own grader. We will report, per the published protocol: highest tier reached, bugs (of 41) that reached each tier, and whether any cell achieved ACE. We will not collapse the ladder into a single pass/fail.
ExploitBench asks researchers not to run reinforcement learning on the set. We will not.
How Panop is run against them
The method is the suites’ method. We do not rewrite the challenges, weaken the defences, or substitute a private flag.
- Isolation. Every target is the published Docker (or ExploitBench) image, started as the maintainers document. Nothing on these hosts is a customer system.
- Briefing. Panop receives what a competitor is supposed to receive: the challenge description and the reachable service (or the published evaluation interface). It does not receive the flag, the solve script,
expected_results/, or the ground-truth vulnerability file. - Flags. On XBOW and Argus the flag is injected at build time, as the READMEs require. A hardcoded string in the agent’s prompt is not a solve.
- Win condition. XBOW and Argus: the injected flag, recovered. ExploitBench: capability bits set by the mechanical grader, not by the agent’s self-report.
- No training on the test. We do not fine-tune, RL, or otherwise fit Panop to these repositories.
- Guardrails stay on. Panop’s production rule — prove access without taking data or leaving a foothold — still applies. A “solve” that requires disabling that rule does not count.
We will also record, as diagnostics rather than as the headline: wall-clock time, steps/turns where the runner exposes them, and (for ExploitBench) cost if the runner reports it. Those numbers will sit next to the scores, not instead of them.
Results
| Suite | Targets | Panop result | Reading |
|---|---|---|---|
| XBOW Validation Benchmarks | 104 flags | Forthcoming | Saturated baseline. A miss would matter; a full score would not, on its own. |
| Argus pentest set | 71 flags | Forthcoming | Discriminating application result. |
| Argus threat-model set | 10 apps | Out of scope | Live exploit validation only, this run. |
| ExploitBench bench-v8 | 41 V8 bugs × 16 capabilities | Forthcoming | Discriminating engine result. Highest tier and ACE count, not a single percentage. |
XBOW
| Level | Challenges | Solved | Rate |
|---|---|---|---|
| 1 (Easy) | — | Forthcoming | — |
| 2 (Medium) | — | Forthcoming | — |
| 3 (Hard) | — | Forthcoming | — |
| Total | 104 | Forthcoming | — |
Argus pentest
| Difficulty | Challenges | Solved | Rate |
|---|---|---|---|
| 1 (Easy) | 2 | Forthcoming | — |
| 2 (Medium) | 27 | Forthcoming | — |
| 3 (Hard) | 42 | Forthcoming | — |
| Total | 71 | Forthcoming | — |
ExploitBench (bench-v8)
| Metric | Panop |
|---|---|
| Highest tier reached | Forthcoming |
| Bugs reaching T5 coverage (of 41) | Forthcoming |
| Bugs reaching T4 reproduction (of 41) | Forthcoming |
| Bugs reaching T3 target primitives (of 41) | Forthcoming |
| Bugs reaching T2 generic primitives (of 41) | Forthcoming |
| Bugs reaching T1 / ACE (of 41) | Forthcoming |
When these tables are filled, the protocol above does not change. If a later run uses a different snapshot of a suite, we will say so on this page rather than silently overwrite.
What this does not prove
These are laboratory targets with a known win condition. They do not replace a scoped engagement against a customer’s estate, and they do not measure whether Panop is safe to run in production — only whether it can demonstrate the exploit the benchmark authors planted.
XBOW, by the maintainers’ own statement, no longer separates strong agents from weak ones. Argus still does, for application pentest work. ExploitBench still does, for engine exploitation, and it measures a climb that most public models currently stall on.
A number without that context is a press release. The context is the point of publishing the protocol first.
The product that this evaluation is supposed to underwrite is autonomous penetration testing: continuous, guardrailed exploitation that turns a severity score into demonstrated access. See also Autonomous Pentest, or book a demo.
Sources
Figures for suite size, difficulty split and scoring rules are the publishers’ own, taken from the repositories and papers linked below. Panop results in the tables are ours, and are empty until the run completes.
- XBOW Engineering — Validation Benchmarks
- Pensar — Argus Validation Benchmarks
- Lee and Brumley — ExploitBench, public leaderboard at exploitbench.ai, paper arXiv:2605.14153
- Panop — Autonomous penetration testing