# VLA safety benchmarks — and what a Provael ASR compares to

> Five benchmarks in the class this project actually competes with, each with a mandatory
> comparability verdict and a mandatory statement of where Provael is WEAKER. These are different
> measurements, not worse ones.

Web page: https://www.provael.com/compare/vla-safety-benchmarks

## The through-line

- Posture: VLA-Arena and SafeVLA-Bench measure safety WITHOUT an adversary — hazard avoidance and post-hoc rollout scoring. Provael measures whether an adversary can move the policy. Different questions; the non-adversarial ones bound a floor that a Provael ASR cannot see.

- Surface: RedVLA attacks the physical scene with the instruction fixed. Provael attacks the instruction with the scene fixed. Complementary halves of the joint space RedVLA itself defines — not a stronger result and a weaker one.

- Evidence: RedVLA and AttackVLA both evaluate on more policies than Provael, and AttackVLA evaluates on real hardware. Provael has ZERO real-robot results. That is a gap, stated as one.

## The benchmarks

### VLA-Arena (arXiv:2512.22539)

- What it is: Capability + safety benchmark with a public leaderboard
- Posture: NON-ADVERSARIAL
- What it perturbs: nothing — hazards are placed in the scene, the instruction is untouched
- Scale: 11 suites, 170 tasks; 5 safety suites, 75 tasks
- Evaluated on: Simulation
- Headline: Cumulative Cost (CC) and Success Rate (SR) per suite
- Comparable to a Provael ASR? NOT DIRECTLY COMPARABLE. The only public VLA leaderboard with a safety axis, which makes it the one place a Provael ASR could be mistaken for a comparable entry. It is not one, and the reason is posture: their suites ask whether a policy is safe by default, Provael asks whether it can be made unsafe. The Provael arm corresponding to their entire safety axis is the BENIGN CONTROL (2/50 on the ten-task run), not any attack family.
- WHERE PROVAEL IS WEAKER: A running public leaderboard with external submissions, against our zero third-party rows. A declarative constraint language (CBDDL) for defining tasks and safety constraints, against our uncalibrated keep-out predicate.
- Source: https://arxiv.org/abs/2512.22539

### SafeVLA-Bench (arXiv:2606.00773)

- What it is: Post-hoc success-safety gap measurement over rollouts
- Posture: NON-ADVERSARIAL
- What it perturbs: nothing — it scores rollouts already produced
- Scale: LIBERO + RoboCasa-365
- Evaluated on: Simulation
- Headline: Succ-But-Unsafe (SBU) and Violation Severity Index (VSI)
- Comparable to a Provael ASR? NOT DIRECTLY COMPARABLE. Post-hoc where Provael is pre-hoc, and nobody causes the failure: it measures the policy’s own behaviour under ordinary instructions. Safety is an STL specification over the trajectory; ours is a boolean from an uncalibrated keep-out predicate. Provael already has a field called succ_but_unsafe that names this benchmark — the shared NAME is exactly why no number is placed beside theirs.
- WHERE PROVAEL IS WEAKER: A declared formal specification (Signal Temporal Logic) where we have an uncalibrated threshold. And their finding bounds ours: high-SR baselines still leave 13–15% unsafe episode rates with no adversary present — a floor a Provael ASR is blind to by construction, because it measures lift over a benign control.
- Source: https://arxiv.org/abs/2606.00773

### RedVLA (arXiv:2604.22591)

- What it is: Physical red-teaming framework + a defense (SimpleVLA-Guard)
- Posture: ADVERSARIAL
- What it perturbs: the physical scene — the instruction is explicitly held fixed
- Scale: 6 VLA models, LIBERO, 10 trials per configuration
- Evaluated on: Simulation
- Headline: ASR up to 95.5% on π₀.₅; 64.9%–95.5% across six models
- Comparable to a Provael ASR? NOT DIRECTLY COMPARABLE. The closest published work to this project, and the cleanest incomparability on the page. RedVLA formalises red teaming as optimisation over the environment–instruction joint space (s′₀, l′), then fixes the instruction (l′ = l) and perturbs only the initial state. Provael does the converse. Same benchmark, same simulator, complementary halves of one formalism — so 95.5% and 88% are orthogonal quantities. Note what this does NOT rest on: both are simulation, so sim-versus-hardware is not the difference here.
- WHERE PROVAEL IS WEAKER: Six policies against our one measured. An optimisation loop over the risk-factor state against our fixed four-template banks. A shipped, evaluated defense against our two stub-validated ones. And a typed physical-safety taxonomy grounded in risk predicates, against a predicate that has never been calibrated on LIBERO.
- Source: https://arxiv.org/abs/2604.22591

### RoboJailBench (arXiv:2605.19328)

- What it is: 18-category harm-outcome benchmark for embodied agents, with a public leaderboard
- Posture: ADVERSARIAL
- What it perturbs: the instruction, against VLM planners
- Scale: RoboVQA, RH20T, NVIDIA PhysicalAI-AV, RJB-Instructions
- Evaluated on: Simulation
- Headline: 94–100% jailbreak ASR in the no-defense setting
- Comparable to a Provael ASR? NOT DIRECTLY COMPARABLE. A different layer of the stack. They evaluate VLM planners — a model that reads a scene and emits a plan — so a success is a harmful sentence. Provael evaluates closed-loop low-level policies emitting motor commands, so a success is a trajectory already outside the envelope. A jailbroken planner has said something; a redirected policy has already moved.
- WHERE PROVAEL IS WEAKER: A public repository with external submissions, against our zero third-party rows — the axis on which this project is weakest, and the one a benchmark is built to be strong on.
- Source: https://arxiv.org/abs/2605.19328

### AttackVLA (arXiv:2511.12149)

- What it is: Unified evaluation framework for adversarial and backdoor attacks on VLAs
- Posture: ADVERSARIAL
- What it perturbs: multiple channels — adversarial and backdoor, including adapted VLM attacks
- Scale: Existing VLA attacks plus attacks adapted from vision-language models
- Evaluated on: Simulation AND real-world robotic settings
- Headline: 58.4% average targeted ASR, reaching 100% on some tasks
- Comparable to a Provael ASR? PARTIALLY. Closest in SHAPE — a unified harness running many attacks and reporting a comparable ASR — and the row that most constrains what this project may claim. Their targeted-attack predicate ("the policy performed the attacker’s specified action sequence") is strictly harder than an envelope breach, so their number is the more demanding one.
- WHERE PROVAEL IS WEAKER: AttackVLA already occupies the position this project is often assumed to have invented — one harness, many attacks, comparable ASR — and it does so WITH real-robot evaluation we do not have. We do not claim to have originated a unified VLA attack harness and we do not claim parity with a benchmark that has been run on hardware.
- Source: https://arxiv.org/abs/2511.12149


## Provael's own row

- What it is: Red-team harness + compliance-evidence emitter
- Posture: ADVERSARIAL
- What it perturbs: the instruction — the scene and the safety envelope are held fixed
- Scale: 1 measured policy (SmolVLA), 10 libero_object tasks, 5 seeds
- Evaluated on: Simulation only — results/hardware/ reads 0
- Headline: ASR 88% (44/50) with a 95% interval and a matched benign control

What is actually different: Not the attacks, and — since AttackVLA — not "a unified harness with a comparable ASR" either. What is left is narrower: a deterministic CPU-only, no-download core so every number is reproducible without a GPU or model weights, and an evidence layer (SARIF, OSCAL, CycloneDX ML-BOM, signed attestation, a Wilson-CI regression gate wired into CI) that a research benchmark has no reason to build.

The VLA-Arena posture contrast has its own page, because that benchmark's leaderboard is the one a
Provael number could actually be mistaken for an entry on: https://www.provael.com/compare/vla-arena

Canonical: https://www.provael.com/compare/vla-safety-benchmarks
