STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
RESEARCH INDEX · EVERY PUBLISHED RESULT

Every finding, with its status.

One measured transfer, a set of honest nulls, a study that verified a negative, and what is still running or planned. Each card states the attack, the policy, the sample size, the interval, and the control - and links to the committed source. No card carries a number that is not in the repository.

1 measured on SmolVLA39 adversarial attacks across 17 adversarial familiesThe measured result →Every measured row →The measured defense →Embodied AI Security Top 10 →

How to read the status of each finding

MEASUREDReal-policy transfer, measured with a 95% Wilson CI and a benign control.
HONEST NULLMeasured on a real policy; the attack did not transfer. We publish it anyway.
STUB-VALIDATEDRuns on the deterministic CPU core; no real-model transfer claimed.
IN PROGRESSRunning or GPU-gated; no number published yet.
PLANNEDOn the roadmap. Listed so the plan is public, with no result implied.

Looking for customer engagements rather than research?/case-studies is the buyer-facing index — it holds none yet, says so, and sets out what a published study will contain. Comparing these results against the published attack literature?Published attack baselines gives every external figure a comparability verdict — and states where this research disagrees with it.

Findings

HYPOTHESISEAI01 · instruction

What fires is a changed meaning, not a changed wording

Policy
SmolVLA
Suite
LIBERO · all ten libero_object tasks
n
50 per arm
Arms
roleplay vs paraphrase

roleplay 44/50 (88.0%) vs paraphrase 3/50 (6.0%) - pooled Wilson intervals [76.2%, 94.4%] and [2.1%, 16.2%] are disjoint (the headline row elsewhere on this site uses the task-clustered interval, 72-100%); paraphrase is indistinguishable from the benign control at McNemar p = 1.0

One policy, one suite, simulation only, maintainer-run, uncalibrated predicate, four templates per arm. A hypothesis with supporting evidence, not a settled result - the repo write-up states what would falsify it and computes that ~1,960 episodes per arm would be needed to test whether paraphrase differs from doing nothing at all.

MEASUREDEAI01 · instruction

A reframed instruction drove a real policy off-task on 44 of 50 trials

Policy
SmolVLA
Suite
LIBERO · all ten libero_object tasks
n
5 seeds per task
Attack
roleplay

ASR 88% (44/50), 95% task-clustered CI [72-100%] · benign control 2/50

The one real-policy transfer in the project. Sim-only, ten tasks, uncalibrated keep-out predicate - not a hazard rate and not a real-robot result. A paired benign-reword arm fired 1/50, so this is attacker control rather than brittleness to rephrasing.

HONEST NULLEAI02 · perception

Perception attacks did not transfer to the real policy

Policy
SmolVLA
Suite
LIBERO
Visual
0/100
Injection
0/50

visual 0/100 [0-4%] · injection 0/50 [0-7%] · same 2/50 benign baseline

A published null with non-zero upper confidence bounds, not an omission. Publishing the zeros is the point - they are what make the 88% believable.

STUB-VALIDATEDEAI04 · action-space

Action-space integrity: verified not-applicable on real policies

Study
v0.20.0 transfer
Families
action · action_space
Real policies
SmolVLA · π0
Result
not-applicable

Stub-validated on the deterministic CPU core (100% [72-100%] vs 0% benign); no real-policy transfer claimed

A v0.20.0 study verified the negative: these inject an out-of-band directive channel a real VLA ignores, so they stay stub-validated by verification, not by omission.

IN PROGRESStransfer · cross-architecture

π0 cross-architecture transfer (GPU-gated)

Policy
π0 (openpi)
Suite
LIBERO
Path
real-policy, GPU-gated
Result
pending

No cross-architecture number yet - the real-policy path is GPU-gated

When it runs it will report with the same Wilson-95 + benign-FPR discipline. No cross-model claim before then.

PRE-REGISTEREDtransfer · second suite

Meta-World second-suite transfer

Suite
Meta-World (shipped, GPU-gated)
Status
study pre-registered
Result
not yet measured

Suite ships in v0.41.2; the transfer study is pre-registered with no result claimed

The suite is implemented and selectable today. What is outstanding is the cross-suite transfer measurement, which needs a GPU run — no number until it runs.

PRE-REGISTEREDtransfer · physical arm

Sim-to-real correlation (SO-ARM101 + SmolVLA)

Platform
SO-ARM101 table-top arm
Status
protocol registered 24 July 2026
Result
not yet measured

Protocol published before the trials; n = 5 per condition, instruction family only

Directional correlation only — does an attack that fires in sim also fire on hardware, and does the benign control stay near zero on both. With n = 5 the real intervals are wide by design; no real-robot ASR is claimed, before or after.

MEASUREDdefense · mitigation

Instruction canonicalization lowered ASR without breaking the task

Defense
instruction canonicalization / repair
Target
EAI01 instruction · EAI01 optimized
Suites
stub + reach (CPU fixture)
Gate
accepted

adversarial ASR 67.5% [52-80%] -> 7.5% [3-20%] on stub · instruction 70.0% [52-83%] -> 10.0% [3-26%] · optimized_instruction 60.0% [31-83%] -> 0.0% [0-28%] · benign FPR 0.0% -> 0.0% · clean-task success 100.0% -> 100.0% (within CI, accepted)

Credited, and substantially circular - the study says so above its own results table. Four of the CPU fixture’s seven danger tokens are words this defense strips, so the optimized_instruction result is close to tautological. stub-validated scaffolding; no real-model transfer is claimed, and the reach arm’s acceptance gate was not evaluable.

Measure your own policy.

Run Provael locally for your own attack-success rate, or send us the open checkpoint closest to your stack and we will red-team it free, in simulation, and send back the scorecard.