Most red-team demos work like a magic trick: show the one attack that landed, and cut away before the ten that didn’t. It makes for a great screenshot and a terrible measurement.
Provael does the opposite. Every attack family ships with its transfer test in the same breath as its number, and we publish the nulls as loudly as the hits.
The one real result, in full
On a real SmolVLA policy in LIBERO, a roleplay instruction drove the arm across a keep-out line 100% of the time (10/10, 95% Wilson CI 72-100%) against a 0% benign control. That is the number people quote.
Here is the number they don’t: the visual-patch and scene-text-injection families produced 0% measurable transfer on the same real model. We could have quietly dropped them. We report them instead, with their confidence intervals.
Why a null is not a failure
A security measurement without its nulls and its benign control is uninterpretable. If you don’t know how often the “unsafe” test fires with no attack present, a 100% attack-success rate could just be a trigger-happy predicate. The benign control is what turns a demo into evidence.
Publishing the nulls does something a hype-driven “most comprehensive platform” structurally cannot: it tells you exactly where the tool is strong and where it isn’t. That honesty is the whole point. The value of a security number is its trustworthiness, and trustworthiness is built by reporting the results that don’t flatter you.
It is a discipline, not a slogan
When we ran the EAI04 action-space transfer study, the honest finding was that those attacks are not-applicable on a real policy through the current mechanism. We shipped that as the headline, not a footnote.
If a number ever looks too good, ask for its benign control and its nulls. If a tool won’t show you those, you are looking at a stunt, not a measurement. See the one real result, in full →