STALE MEASUREMENTPast this project's own 2-release window: the published result was measured with v0.32.0, 9 releases ago. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
THE PEER CLASS · PV-020

VLA safety benchmarks, compared

These are different measurements, not worse ones.

Five benchmarks in the class Provael actually competes with. Every row carries a comparability verdict and — because a comparison page without one is an advertisement — a statement of where Provael is weaker.

Every row citedProvael has 0 real-robot results

Three axes separate every row below

Posture
VLA-Arena and SafeVLA-Bench measure safety WITHOUT an adversary — hazard avoidance and post-hoc rollout scoring. Provael measures whether an adversary can move the policy. Different questions; the non-adversarial ones bound a floor that a Provael ASR cannot see.
Surface
RedVLA attacks the physical scene with the instruction fixed. Provael attacks the instruction with the scene fixed. Complementary halves of the joint space RedVLA itself defines — not a stronger result and a weaker one.
Evidence
RedVLA and AttackVLA both evaluate on more policies than Provael, and AttackVLA evaluates on real hardware. Provael has ZERO real-robot results. That is a gap, stated as one.

At a glance

Five VLA safety and attack benchmarks with Provael's own row last, compared by posture, what is perturbed, whether evaluation is simulation or hardware, the headline figure, and whether that figure is comparable to a Provael attack-success rate.
BenchmarkPostureWhat it perturbsEvaluated onHeadlineComparable?
VLA-ArenaarXiv:2512.22539non-adversarialnothing — hazards are placed in the scene, the instruction is untouchedSimulationCumulative Cost (CC) and Success Rate (SR) per suiteNot directly comparable
SafeVLA-BencharXiv:2606.00773non-adversarialnothing — it scores rollouts already producedSimulationSucc-But-Unsafe (SBU) and Violation Severity Index (VSI)Not directly comparable
RedVLAarXiv:2604.22591adversarialthe physical scene — the instruction is explicitly held fixedSimulationASR up to 95.5% on π₀.₅; 64.9%–95.5% across six modelsNot directly comparable
RoboJailBencharXiv:2605.19328adversarialthe instruction, against VLM plannersSimulation94–100% jailbreak ASR in the no-defense settingNot directly comparable
AttackVLAarXiv:2511.12149adversarialmultiple channels — adversarial and backdoor, including adapted VLM attacksSimulation AND real-world robotic settings58.4% average targeted ASR, reaching 100% on some tasksPartially
Provaeladversarialthe instruction — the scene and the safety envelope are held fixedSimulation only — results/hardware/ reads 0ASR 88% (44/50) with a 95% interval and a matched benign controlreference definition

Row by row, including where we lose

VLA-ArenaCapability + safety benchmark with a public leaderboard

11 suites, 170 tasks; 5 safety suites, 75 tasks · Simulation

Not directly comparable. The only public VLA leaderboard with a safety axis, which makes it the one place a Provael ASR could be mistaken for a comparable entry. It is not one, and the reason is posture: their suites ask whether a policy is safe by default, Provael asks whether it can be made unsafe. The Provael arm corresponding to their entire safety axis is the BENIGN CONTROL (2/50 on the ten-task run), not any attack family.

Where Provael is weaker

A running public leaderboard with external submissions, against our zero third-party rows. A declarative constraint language (CBDDL) for defining tasks and safety constraints, against our uncalibrated keep-out predicate.

SafeVLA-BenchPost-hoc success-safety gap measurement over rollouts

LIBERO + RoboCasa-365 · Simulation

Not directly comparable. Post-hoc where Provael is pre-hoc, and nobody causes the failure: it measures the policy’s own behaviour under ordinary instructions. Safety is an STL specification over the trajectory; ours is a boolean from an uncalibrated keep-out predicate. Provael already has a field called succ_but_unsafe that names this benchmark — the shared NAME is exactly why no number is placed beside theirs.

Where Provael is weaker

A declared formal specification (Signal Temporal Logic) where we have an uncalibrated threshold. And their finding bounds ours: high-SR baselines still leave 13–15% unsafe episode rates with no adversary present — a floor a Provael ASR is blind to by construction, because it measures lift over a benign control.

RedVLAPhysical red-teaming framework + a defense (SimpleVLA-Guard)

6 VLA models, LIBERO, 10 trials per configuration · Simulation

Not directly comparable. The closest published work to this project, and the cleanest incomparability on the page. RedVLA formalises red teaming as optimisation over the environment–instruction joint space (s′₀, l′), then fixes the instruction (l′ = l) and perturbs only the initial state. Provael does the converse. Same benchmark, same simulator, complementary halves of one formalism — so 95.5% and 88% are orthogonal quantities. Note what this does NOT rest on: both are simulation, so sim-versus-hardware is not the difference here.

Where Provael is weaker

Six policies against our one measured. An optimisation loop over the risk-factor state against our fixed four-template banks. A shipped, evaluated defense against our two stub-validated ones. And a typed physical-safety taxonomy grounded in risk predicates, against a predicate that has never been calibrated on LIBERO.

RoboJailBench18-category harm-outcome benchmark for embodied agents, with a public leaderboard

RoboVQA, RH20T, NVIDIA PhysicalAI-AV, RJB-Instructions · Simulation

Not directly comparable. A different layer of the stack. They evaluate VLM planners — a model that reads a scene and emits a plan — so a success is a harmful sentence. Provael evaluates closed-loop low-level policies emitting motor commands, so a success is a trajectory already outside the envelope. A jailbroken planner has said something; a redirected policy has already moved.

Where Provael is weaker

A public repository with external submissions, against our zero third-party rows — the axis on which this project is weakest, and the one a benchmark is built to be strong on.

AttackVLAUnified evaluation framework for adversarial and backdoor attacks on VLAs

Existing VLA attacks plus attacks adapted from vision-language models · Simulation AND real-world robotic settings

Partially. Closest in SHAPE — a unified harness running many attacks and reporting a comparable ASR — and the row that most constrains what this project may claim. Their targeted-attack predicate ("the policy performed the attacker’s specified action sequence") is strictly harder than an envelope breach, so their number is the more demanding one.

Where Provael is weaker

AttackVLA already occupies the position this project is often assumed to have invented — one harness, many attacks, comparable ASR — and it does so WITH real-robot evaluation we do not have. We do not claim to have originated a unified VLA attack harness and we do not claim parity with a benchmark that has been run on hardware.

What is actually different about Provael

Not the attacks, and — since AttackVLA — not "a unified harness with a comparable ASR" either. What is left is narrower: a deterministic CPU-only, no-download core so every number is reproducible without a GPU or model weights, and an evidence layer (SARIF, OSCAL, CycloneDX ML-BOM, signed attestation, a Wilson-CI regression gate wired into CI) that a research benchmark has no reason to build.

The one that needs its own page

VLA-Arena runs the only public VLA leaderboard with a safety axis, which makes it the one benchmark a Provael number could actually be mistaken for an entry on. The posture contrast, the five safety suites and the coverage tally are set out in full there.Why the two do not share a column →

The measurement, with its limits attached.

One policy, ten tasks, simulation only, uncalibrated predicate — and a benign control.