STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
THE ADJACENT GIANT · PV-029

Provael and Gray Swan are not the same kind of thing

They run a competition. This runs a measurement.

Gray Swan AI operates the largest AI red-teaming arena there is, and on their own published figures it is bigger than this project by three to five orders of magnitude. That is the first thing this page should say, so here it is at the top rather than in a footnote. What separates the two is not size: it is what comes out of them.

Their figures, from their own pageProvael is smaller on every measure of scale

Their scale, in their numbers

Read from the Arena’s own About page on 6 September 2026. These are Gray Swan’s published figures, reproduced rather than restated.

$490K+
in rewards distributed
4M+
attack attempts submitted
130K+
successful breaks
13K+
community members
100+
participants placed in paid red-teaming roles

Seven axes, including the four where Provael loses

Gray Swan and Provael compared on seven axes: who attacks, scale, what is under attack, what counts as success, what comes out, reproducibility, and who is behind each project. The last column states which of the two is ahead on that axis, or that the axis is one where the two answer different questions.
AxisGray SwanProvaelWhich is ahead
Who attacksPeople. 13K+ community members competing for prize pools from $20,000 to $170,000+, submitting through a chat interface. Their About page notes "no coding required".A program. 39 registered adversarial attacks resolved by name and run by a fixed harness, plus a benign control arm. Nobody is being creative at run time.Gray Swan is ahead
Scale4M+ attack attempts and 130K+ successful breaks, on their own figures.29 committed measurements, of which one is a real-model suite result: one policy, ten tasks, 50 episodes per attack arm.Gray Swan is ahead
What is under attackFrontier language models and software agents with tools, memory and autonomous capabilities. Their published challenge types are chat, vision, agents, reasoning, indirect prompt injection and cyber/CTF.A vision-language-action policy driving a simulated robot. The attack surface is the instruction and the observation; the thing scored is the trajectory.Different questions
What counts as successA "break" — a submission that triggers a target behaviour, judged by an AI panel with human appeals on some challenges.A keep-out envelope violation, scored by a predicate over the simulated end-effector pose. The predicate is uncalibrated and the page carrying the number says so.Different questions
What comes outBreaks, held private until 30 days after a challenge ends, and research findings shared with the model developers.An attack-success rate with a 95% Wilson confidence interval against a benign control, plus SARIF, OSCAL assessment-results and a CycloneDX ML-BOM. Committed to a public repository at the moment it is produced.Provael is ahead
ReproducibilityNot the goal. A competition is a search by thousands of people and cannot be re-run to the same answer, which is what makes it good at finding the unexpected.The whole design. A run is a pure function of its config and seed, so the same command yields a byte-identical report, and an execution manifest binds the run to its provenance.Provael is ahead
Who is behind itA company, with UK AISI, OpenAI, Anthropic, Amazon and Meta named as Frontier Lab sponsors of the Arena.One maintainer, Apache-2.0, no funding.Gray Swan is ahead

What the Arena publishes as its challenge types

Six, listed on their About page on 6 September 2026:

  • Chat
  • Image / Vision
  • Agents (tools, memory, autonomous capabilities)
  • Reasoning
  • Indirect prompt injection (instructions hidden in data the AI processes)
  • Cyber / CTF

None of them is an embodied or physical category, and that is the whole of the observation. It is a statement about a page on a date, not about what Gray Swan can do, intends to do, or has decided against — they have said nothing of the kind, and a company running competitions against frontier models could add a robot target whenever it chose to. Read their About page before relying on this.

Where this leaves Provael weaker

Thirteen thousand people inventing attacks will find things a registry of thirty-nine cannot. That is not a gap a fixed harness closes by adding attacks, because the mechanism is different: a competition searches, and a harness re-runs. The honest read is that a crowd finds the unknown attack and a harness re-measures the known one on every commit, and a team that needs the first should not buy the second instead.

A break and a rate answer different questions

This is the difference worth carrying away, and it survives the scale gap entirely.

A break says it can happen

Somebody got the system to do the thing. That is a proof of existence, it is exactly what you want before shipping, and no amount of averaging substitutes for it. It is also, by construction, not a rate: one break out of an unrecorded number of attempts by an unknown number of people is not a denominator, and Gray Swan does not present it as one.

A rate says how often, with an interval and a control

A fixed attack, a fixed policy, a fixed number of episodes, and the same measurement again next week against the same benign control. That is what an assessor cites and what a regression gate compares against, and it is worth nothing at all if the attack in it is one nobody would have tried. A rate needs a search to have happened first.

The measurement, with its limits attached.

One policy, ten tasks, simulation only, uncalibrated predicate — and a benign control.