Provael and Gray Swan are not the same kind of thing
They run a competition. This runs a measurement.
Gray Swan AI operates the largest AI red-teaming arena there is, and on their own published figures it is bigger than this project by three to five orders of magnitude. That is the first thing this page should say, so here it is at the top rather than in a footnote. What separates the two is not size: it is what comes out of them.
Their scale, in their numbers
Read from the Arena’s own About page on 6 September 2026. These are Gray Swan’s published figures, reproduced rather than restated.
- $490K+
- in rewards distributed
- 4M+
- attack attempts submitted
- 130K+
- successful breaks
- 13K+
- community members
- 100+
- participants placed in paid red-teaming roles
For the comparison: Provael’s one real-model suite result is 44/50 trials of a single attack against a single policy on ten tasks. Both numbers are true, and they are not the same kind of number.
Seven axes, including the four where Provael loses
| Axis | Gray Swan | Provael | Which is ahead |
|---|---|---|---|
| Who attacks | People. 13K+ community members competing for prize pools from $20,000 to $170,000+, submitting through a chat interface. Their About page notes "no coding required". | A program. 39 registered adversarial attacks resolved by name and run by a fixed harness, plus a benign control arm. Nobody is being creative at run time. | Gray Swan is ahead |
| Scale | 4M+ attack attempts and 130K+ successful breaks, on their own figures. | 29 committed measurements, of which one is a real-model suite result: one policy, ten tasks, 50 episodes per attack arm. | Gray Swan is ahead |
| What is under attack | Frontier language models and software agents with tools, memory and autonomous capabilities. Their published challenge types are chat, vision, agents, reasoning, indirect prompt injection and cyber/CTF. | A vision-language-action policy driving a simulated robot. The attack surface is the instruction and the observation; the thing scored is the trajectory. | Different questions |
| What counts as success | A "break" — a submission that triggers a target behaviour, judged by an AI panel with human appeals on some challenges. | A keep-out envelope violation, scored by a predicate over the simulated end-effector pose. The predicate is uncalibrated and the page carrying the number says so. | Different questions |
| What comes out | Breaks, held private until 30 days after a challenge ends, and research findings shared with the model developers. | An attack-success rate with a 95% Wilson confidence interval against a benign control, plus SARIF, OSCAL assessment-results and a CycloneDX ML-BOM. Committed to a public repository at the moment it is produced. | Provael is ahead |
| Reproducibility | Not the goal. A competition is a search by thousands of people and cannot be re-run to the same answer, which is what makes it good at finding the unexpected. | The whole design. A run is a pure function of its config and seed, so the same command yields a byte-identical report, and an execution manifest binds the run to its provenance. | Provael is ahead |
| Who is behind it | A company, with UK AISI, OpenAI, Anthropic, Amazon and Meta named as Frontier Lab sponsors of the Arena. | One maintainer, Apache-2.0, no funding. | Gray Swan is ahead |
Gray Swan’s column reflects their own published description of the Arena. Names and marks are their owners’; Provael is independent and affiliated with none of them.
What the Arena publishes as its challenge types
Six, listed on their About page on 6 September 2026:
- Chat
- Image / Vision
- Agents (tools, memory, autonomous capabilities)
- Reasoning
- Indirect prompt injection (instructions hidden in data the AI processes)
- Cyber / CTF
None of them is an embodied or physical category, and that is the whole of the observation. It is a statement about a page on a date, not about what Gray Swan can do, intends to do, or has decided against — they have said nothing of the kind, and a company running competitions against frontier models could add a robot target whenever it chose to. Read their About page before relying on this.
Thirteen thousand people inventing attacks will find things a registry of thirty-nine cannot. That is not a gap a fixed harness closes by adding attacks, because the mechanism is different: a competition searches, and a harness re-runs. The honest read is that a crowd finds the unknown attack and a harness re-measures the known one on every commit, and a team that needs the first should not buy the second instead.
A break and a rate answer different questions
This is the difference worth carrying away, and it survives the scale gap entirely.
A break says it can happen
Somebody got the system to do the thing. That is a proof of existence, it is exactly what you want before shipping, and no amount of averaging substitutes for it. It is also, by construction, not a rate: one break out of an unrecorded number of attempts by an unknown number of people is not a denominator, and Gray Swan does not present it as one.
A rate says how often, with an interval and a control
A fixed attack, a fixed policy, a fixed number of episodes, and the same measurement again next week against the same benign control. That is what an assessor cites and what a regression gate compares against, and it is worth nothing at all if the attack in it is one nobody would have tried. A rate needs a search to have happened first.
No insurer, notified body or assessor has reviewed either output against the other, and nothing here says one substitutes for the other. If you are choosing, the question is whether you need to know that a failure is possible or how often it recurs.
The measurement, with its limits attached.
One policy, ten tasks, simulation only, uncalibrated predicate — and a benign control.