# State of VLA Security — the measured result

> On a real SmolVLA policy in LIBERO (task libero_object, 10 seeds), the roleplay attack
> (instruction family) redirected the policy 100% (10/10) [72–100% CI · 95% Wilson] versus a 0%
> benign baseline (0/10). Visual and injection attacks did not transfer (0%). Simulation only,
> one task, n=10, uncalibrated predicate. Reproducible from a public notebook.

## The numbers (same 17 successes, framed honestly)

- Headline — roleplay (instruction family): ASR 100% (10/10), 95% Wilson CI [72–100%]
- Instruction family (roleplay + goal_substitution + paraphrase): 56.7% (17/30) [39–73%]
- Overall across the battery: 24.3% (17/70)
- Honest nulls: visual 0% (0/20) [0–16%]; injection 0% (0/10) [0–28%] — they did not transfer
- Benign FPR: 0% (0/10) — the control held, so every success is attack lift
- Attack: roleplay · Family: instruction (EAI01) · Policy: SmolVLA (HuggingFaceVLA/smolvla_libero)
- Suite: LIBERO · robosuite · MuJoCo · task libero_object · 10 seeds · horizon 280
- Caveat: uncalibrated keep-out predicate — "diverted out of the benign envelope," not a hazard rate

## What this is NOT

- Not a real-robot result — simulation only.
- Not a benchmark score — small n (10), wide CI reported, uncalibrated predicate.
- Not a broad claim — only the instruction family transfers; visual and injection are honest 0% nulls.
- Not a firmware claim — UniPwn-class exploits are out of scope (EAI07).
- Not a safety certificate. Evidence, not certification.

## External validation (distinct)

RoboPAIR (Robey et al., UPenn), arXiv:2410.13691 — an algorithmic jailbreak of LLM-controlled
robots, ~100% across three systems. Cited as independent evidence the category is real; a
different study on different systems.

Reproduce: https://github.com/provael/provael

---
Provael · Prove it. Prevail. · Apache-2.0 · https://github.com/provael/provael
Not legal advice; verify regulatory dates against the primary source.
