# Provael and Gray Swan are not the same kind of thing

> Gray Swan AI operates the largest AI red-teaming arena there is, and on their own
> published figures it is bigger than this project by three to five orders of magnitude. That is
> stated here rather than buried. What separates the two is not size: a competition produces a
> BREAK, and this produces a RATE with a confidence interval and a benign control.

Web page: https://www.provael.com/compare/gray-swan

## Their scale, in their numbers

Read from the Arena's own About page on 6 September 2026. Gray Swan's published figures,
reproduced rather than restated.

- $490K+ in rewards distributed
- 4M+ attack attempts submitted
- 130K+ successful breaks
- 13K+ community members
- 100+ participants placed in paid red-teaming roles

For the comparison: Provael's one real-model suite result is 44/50 trials of a single
attack against a single policy on ten tasks.

## Seven axes, including the four where Provael loses

### Who attacks

- **Gray Swan:** People. 13K+ community members competing for prize pools from $20,000 to $170,000+, submitting through a chat interface. Their About page notes "no coding required".
- **Provael:** A program. 39 registered adversarial attacks resolved by name and run by a fixed harness, plus a benign control arm. Nobody is being creative at run time.
- **Which is ahead:** Gray Swan is ahead

### Scale

- **Gray Swan:** 4M+ attack attempts and 130K+ successful breaks, on their own figures.
- **Provael:** 29 committed measurements, of which one is a real-model suite result: one policy, ten tasks, 50 episodes per attack arm.
- **Which is ahead:** Gray Swan is ahead

### What is under attack

- **Gray Swan:** Frontier language models and software agents with tools, memory and autonomous capabilities. Their published challenge types are chat, vision, agents, reasoning, indirect prompt injection and cyber/CTF.
- **Provael:** A vision-language-action policy driving a simulated robot. The attack surface is the instruction and the observation; the thing scored is the trajectory.
- **Which is ahead:** Different questions

### What counts as success

- **Gray Swan:** A "break" — a submission that triggers a target behaviour, judged by an AI panel with human appeals on some challenges.
- **Provael:** A keep-out envelope violation, scored by a predicate over the simulated end-effector pose. The predicate is uncalibrated and the page carrying the number says so.
- **Which is ahead:** Different questions

### What comes out

- **Gray Swan:** Breaks, held private until 30 days after a challenge ends, and research findings shared with the model developers.
- **Provael:** An attack-success rate with a 95% Wilson confidence interval against a benign control, plus SARIF, OSCAL assessment-results and a CycloneDX ML-BOM. Committed to a public repository at the moment it is produced.
- **Which is ahead:** Provael is ahead

### Reproducibility

- **Gray Swan:** Not the goal. A competition is a search by thousands of people and cannot be re-run to the same answer, which is what makes it good at finding the unexpected.
- **Provael:** The whole design. A run is a pure function of its config and seed, so the same command yields a byte-identical report, and an execution manifest binds the run to its provenance.
- **Which is ahead:** Provael is ahead

### Who is behind it

- **Gray Swan:** A company, with UK AISI, OpenAI, Anthropic, Amazon and Meta named as Frontier Lab sponsors of the Arena.
- **Provael:** One maintainer, Apache-2.0, no funding.
- **Which is ahead:** Gray Swan is ahead


## What the Arena publishes as its challenge types

- Chat
- Image / Vision
- Agents (tools, memory, autonomous capabilities)
- Reasoning
- Indirect prompt injection (instructions hidden in data the AI processes)
- Cyber / CTF

None of them is an embodied or physical category, and that is the whole of the observation. It is a
statement about a page on a date, not about what Gray Swan can do, intends to do, or has decided
against. They have said nothing of the kind, and a company running competitions against frontier
models could add a robot target whenever it chose to. Read their About page before relying on this:
https://app.grayswan.ai/arena/about

## Where this leaves Provael weaker

Thirteen thousand people inventing attacks will find things a registry of thirty-nine cannot. That
is not a gap a fixed harness closes by adding attacks, because the mechanism is different: a
competition searches, and a harness re-runs. A crowd finds the unknown attack; a harness re-measures
the known one on every commit. A team that needs the first should not buy the second instead.

## A break and a rate answer different questions

A BREAK says it can happen. Somebody got the system to do the thing — a proof of existence, exactly
what you want before shipping, and not a rate: one break out of an unrecorded number of attempts by
an unknown number of people has no denominator, and Gray Swan does not present it as one.

A RATE says how often, with an interval and a control. A fixed attack, a fixed policy, a fixed
number of episodes, and the same measurement again next week against the same benign control. That
is what an assessor cites and what a regression gate compares against — and it is worth nothing if
the attack in it is one nobody would have tried. A rate needs a search to have happened first.

No insurer, notified body or assessor has reviewed either output against the other, and nothing here
says one substitutes for the other. Names and marks are their owners'; Provael is independent and
affiliated with none of them.

Canonical: https://www.provael.com/compare/gray-swan
