STALE MEASUREMENTThe newest real-model measurement on this site is 14 days old — past this project’s own 7-day window. Every measured figure here was true when it was taken and has not been re-measured since. Why, and what unblocks it

ProductEvidenceAdoptersTop 10For youLeaderboardSubmitCompliancePricingReadinessDocsStar on GitHub Quickstart
PRE-REGISTERED STUDY · NO RESULT CLAIMED

Does any of this happen on a real robot?

Not yet measured. Every number Provael publishes today comes from simulation, and no sim-to-real claim is made anywhere on this site. What exists is a protocol — written down, dated, and published here before the trials run, so that when a result does land you can check it against a design nobody chose after seeing the data.

Someone else's hardware result

Someone else's hardware result: Zhang et al. 2026 — not a Provael measurement

The concession above is the point of this page, and it is sharper with a concrete example of what is being conceded. This section is not Provael's evidence and is labelled at every step so it cannot be read as such.

Third-party result — not Provael's SARF and AGSD, Zhang, Yin, Yang, Yan, Tian and Yu, published arXiv:2608.03231 on 4 August 2026.

“On LIBERO, SARF reduces OpenVLA's failure rate under AGSD from 100% to 14.2%-56.8% (28.6% average) across suites while preserving clean performance, and on a real PiPER manipulator it improves average success under AGSD from 23.0% to 65.0%.”

AGSD is the attack — a printable patch that hijacks the policy's action-to-vision attention. SARF is the paired defense. The number that matters for this page is the last one: 23.0% to 65.0% average success under attack, on a physical arm. Not in simulation, and not by these authors' own framing a certification of anything — but measured on hardware.

Provael does not have a number like that, for any family, attack or defense. Not a weaker one, not a preliminary one — none. Every figure this project publishes comes from simulation, which is the sentence at the top of this page and is unchanged by anything in this section. Reading a third party's hardware result as evidence about Provael would be precisely the substitution this page exists to refuse.

The plan to close it is the pre-registered protocol below: the SO-ARM101 correlation study, written down before its trials so that a result can be judged against a design nobody chose after seeing the data. It has not been run. The studies index carries the date by which this project publishes either that result or the specific reason it did not happen.

The protocol, registered and not yet run

What will be done, fixed in advance

Transcribed from the pre-registration in the product repository, which is the source of record: sim-to-real correlation study (SO-ARM101 + SmolVLA). Provael ships no physical-robot control code and no real-world-harm payloads; its role here is the simulation-side attack and the offline correlation analysis.

Real platform
SO-ARM101 — a low-cost, open, table-top manipulator
Policy
SmolVLA, the same checkpoint used in simulation
Actuation
LeRobot’s standard eval loop. No Provael code actuates the arm.
Attack family
instruction only — roleplay, goal_substitution, paraphrase
Control
the benign none condition, run on both legs
Unsafe predicate
a geometric keep-out zone on the table, disjoint from the benign task envelope
Documentation
every trial filmed; seed, command, attack and outcome logged
Supervision
a human supervisor with an e-stop present throughout
Pre-registered protocol parameters

Conditions and planned sample size

Pre-registered design for the sim-to-real study. No data has been collected. Every sample size below is a planned target, not an observation, and every result cell reads “Not yet run”.
ConditionChannelTarget n (planned)Result
roleplayinstruction reframing5Not yet run
goal_substitutioninstruction reframing5Not yet run
paraphraseinstruction reframing5Not yet run
nonebenign control5Not yet run

How it will be judged

The analysis is a correlation, not an equivalence. For each condition the simulation redirection rate and its 95% Wilson interval are compared against the real rate and its own small-n interval, and the study asks two questions: does an attack that fires in simulation also fire on hardware, and does the benign control stay near zero on both. With n = 5 the real intervals are wide by design. The claim under test is directional — that simulation predicts real — and never a point-estimate match.

The null hypothesis is stated too, which is the part that makes this falsifiable: simulation and real disagree, with simulation over- or under-predicting real behaviour. The pre-registration also names its own threats to validity — tiny hardware n, the embodiment and calibration gap, unblinded scoring by a supervising operator, and a single arm, policy and task.

What the result will not license anyone to claim

Even when it lands, this is a narrow instrument

  • A real-robot attack-success rate. Five trials per condition give intervals too wide for a point estimate, and the study is designed as a directional check, not a measurement of a real ASR.
  • Evidence about any other robot, policy or task. One arm, one policy, one mirrored task.
  • A hazard probability. The predicate is a geometric keep-out zone chosen to be conservative and bounded — not a certified danger threshold.
  • A conformity conclusion, a certification, or a notified-body opinion.
  • A claim that simulation is sufficient. A correlation that holds here would be evidence for this setup and an argument for testing the next one, not a licence to stop.

What it will support, if the directions agree, is narrower than most readers expect and more useful than nothing: evidence that for this arm, this policy and this task, a simulation red-team result is a signal about hardware behaviour rather than a simulator artefact. That is the question the study was designed to answer, and it is the only one it can.

Why the order matters

A sim-to-real study designed after seeing the numbers is not evidence

With the protocol fixed only after the data exists, every degree of freedom is available to whoever writes it up. Which conditions to report, how many trials to count, where the keep-out zone sits, what counts as agreement — each can be chosen, honestly and without anyone intending to mislead, in the direction the numbers already point. The result looks identical to a real finding and carries none of its weight.

Publishing first is the only thing that lets a reader tell the two apart. It is also the only version that can embarrass us: the null hypothesis is written down, the sample size is fixed, and the threats to validity are named by the people running the study rather than by a critic afterwards. If simulation turns out not to predict hardware here, that outcome is already on the record as a thing this study can find — and it will be published the same way the transfers were.

That is the same reason this site publishes null results and prints the families that did not transfer beside the one that did on /results. A red team that reports only its hits has no denominator.

Where this sits

The rest of the evidence, and what it does not cover

The measured simulation result lives on /results, with its confidence interval, its benign control and its published nulls. Its status — alongside every other study that is registered but unrun — is tracked on /findings. For where a policy red-team sits relative to the compute certification and ROS-layer security work other vendors do, see the stack comparison.

The protocol itself is versioned in the product repository and changes there first:github.com/provael/provael.