<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Provael — findings</title><description>Measured attack-success results and honest nulls from red-teaming VLA robot policies in simulation, each with its n, 95% Wilson interval and benign control.</description><link>https://www.provael.com/</link><language>en</language><copyright>Content CC BY-SA 4.0 · Provael</copyright><item><title>HYPOTHESIS: What fires is a changed meaning, not a changed wording</title><link>https://www.provael.com/findings#eai01-instruction</link><guid isPermaLink="true">https://www.provael.com/findings#eai01-instruction</guid><description>roleplay 44/50 (88.0%) [76.2%, 94.4%] vs paraphrase 3/50 (6.0%) [2.1%, 16.2%] - intervals disjoint; paraphrase is indistinguishable from the benign control at McNemar p = 1.0</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;EAI01 · instruction&lt;/strong&gt; — HYPOTHESIS&lt;/p&gt;&lt;p&gt;roleplay 44/50 (88.0%) [76.2%, 94.4%] vs paraphrase 3/50 (6.0%) [2.1%, 16.2%] - intervals disjoint; paraphrase is indistinguishable from the benign control at McNemar p = 1.0&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Policy: SmolVLA&lt;/li&gt;&lt;li&gt;Suite: LIBERO · all ten libero_object tasks&lt;/li&gt;&lt;li&gt;n: 50 per arm&lt;/li&gt;&lt;li&gt;Arms: roleplay vs paraphrase&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;One policy, one suite, simulation only, maintainer-run, uncalibrated predicate, four templates per arm. A hypothesis with supporting evidence, not a settled result - the repo write-up states what would falsify it and computes that ~1,960 episodes per arm would be needed to test whether paraphrase differs from doing nothing at all.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael/blob/main/docs/findings/semantic-vs-mechanical-instruction-attacks.md&quot;&gt;The full write-up (repo) ↗&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/results/&quot;&gt;The measured result&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>hypothesis</category></item><item><title>MEASURED: A reframed instruction drove a real policy off-task on 44 of 50 trials</title><link>https://www.provael.com/findings#eai01-instruction</link><guid isPermaLink="true">https://www.provael.com/findings#eai01-instruction</guid><description>ASR 88% (44/50), 95% task-clustered CI [72-100%] · benign control 2/50</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;EAI01 · instruction&lt;/strong&gt; — MEASURED&lt;/p&gt;&lt;p&gt;ASR 88% (44/50), 95% task-clustered CI [72-100%] · benign control 2/50&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Policy: SmolVLA&lt;/li&gt;&lt;li&gt;Suite: LIBERO · all ten libero_object tasks&lt;/li&gt;&lt;li&gt;n: 5 seeds per task&lt;/li&gt;&lt;li&gt;Attack: roleplay&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;The one real-policy transfer in the project. Sim-only, ten tasks, uncalibrated keep-out predicate - not a hazard rate and not a real-robot result. A paired benign-reword arm fired 1/50, so this is attacker control rather than brittleness to rephrasing.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/results/&quot;&gt;The measured result&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael/tree/ae828bf8c1bf8ba09c79058553449a6f4dd50a25/results/smolvla_libero_object_suite&quot;&gt;Immutable report ↗&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>measured</category></item><item><title>HONEST NULL: Perception attacks did not transfer to the real policy</title><link>https://www.provael.com/findings#eai02-perception</link><guid isPermaLink="true">https://www.provael.com/findings#eai02-perception</guid><description>visual 0/100 [0-4%] · injection 0/50 [0-7%] · same 0/10 benign baseline</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;EAI02 · perception&lt;/strong&gt; — HONEST NULL&lt;/p&gt;&lt;p&gt;visual 0/100 [0-4%] · injection 0/50 [0-7%] · same 0/10 benign baseline&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Policy: SmolVLA&lt;/li&gt;&lt;li&gt;Suite: LIBERO&lt;/li&gt;&lt;li&gt;Visual: 0/100&lt;/li&gt;&lt;li&gt;Injection: 0/50&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;A published null with non-zero upper confidence bounds, not an omission. Publishing the zeros is the point - they are what make the 100% believable.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/eai-top-10/eai02/&quot;&gt;How we test perception&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/results/&quot;&gt;Full breakdown&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>honest-null</category></item><item><title>STUB-VALIDATED: Action-space integrity: verified not-applicable on real policies</title><link>https://www.provael.com/findings#eai04-action-space</link><guid isPermaLink="true">https://www.provael.com/findings#eai04-action-space</guid><description>Stub-validated on the deterministic CPU core (100% [72-100%] vs 0% benign); no real-policy transfer claimed</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;EAI04 · action-space&lt;/strong&gt; — STUB-VALIDATED&lt;/p&gt;&lt;p&gt;Stub-validated on the deterministic CPU core (100% [72-100%] vs 0% benign); no real-policy transfer claimed&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Study: v0.20.0 transfer&lt;/li&gt;&lt;li&gt;Families: action · action_space&lt;/li&gt;&lt;li&gt;Real policies: SmolVLA · π0&lt;/li&gt;&lt;li&gt;Result: not-applicable&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;A v0.20.0 study verified the negative: these inject an out-of-band directive channel a real VLA ignores, so they stay stub-validated by verification, not by omission.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael/blob/main/docs/studies/eai04-action-space-transfer.md&quot;&gt;The EAI04 action-space transfer study ↗&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>stub-validated</category></item><item><title>IN PROGRESS: π0 cross-architecture transfer (GPU-gated)</title><link>https://www.provael.com/findings#transfer-cross-architecture</link><guid isPermaLink="true">https://www.provael.com/findings#transfer-cross-architecture</guid><description>No cross-architecture number yet - the real-policy path is GPU-gated</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;transfer · cross-architecture&lt;/strong&gt; — IN PROGRESS&lt;/p&gt;&lt;p&gt;No cross-architecture number yet - the real-policy path is GPU-gated&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Policy: π0 (openpi)&lt;/li&gt;&lt;li&gt;Suite: LIBERO&lt;/li&gt;&lt;li&gt;Path: real-policy, GPU-gated&lt;/li&gt;&lt;li&gt;Result: pending&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;When it runs it will report with the same Wilson-95 + benign-FPR discipline. No cross-model claim before then.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael&quot;&gt;Follow on GitHub ↗&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>in-progress</category></item><item><title>PRE-REGISTERED: Meta-World second-suite transfer</title><link>https://www.provael.com/findings#transfer-second-suite</link><guid isPermaLink="true">https://www.provael.com/findings#transfer-second-suite</guid><description>Suite ships in v0.41.1; the transfer study is pre-registered with no result claimed</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;transfer · second suite&lt;/strong&gt; — PRE-REGISTERED&lt;/p&gt;&lt;p&gt;Suite ships in v0.41.1; the transfer study is pre-registered with no result claimed&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Suite: Meta-World (shipped, GPU-gated)&lt;/li&gt;&lt;li&gt;Status: study pre-registered&lt;/li&gt;&lt;li&gt;Result: not yet measured&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;The suite is implemented and selectable today. What is outstanding is the cross-suite transfer measurement, which needs a GPU run — no number until it runs.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael&quot;&gt;Roadmap ↗&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>pre-registered</category></item><item><title>PRE-REGISTERED: Sim-to-real correlation (SO-ARM101 + SmolVLA)</title><link>https://www.provael.com/findings#transfer-physical-arm</link><guid isPermaLink="true">https://www.provael.com/findings#transfer-physical-arm</guid><description>Protocol published before the trials; n = 5 per condition, instruction family only</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;transfer · physical arm&lt;/strong&gt; — PRE-REGISTERED&lt;/p&gt;&lt;p&gt;Protocol published before the trials; n = 5 per condition, instruction family only&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Platform: SO-ARM101 table-top arm&lt;/li&gt;&lt;li&gt;Status: protocol registered 24 July 2026&lt;/li&gt;&lt;li&gt;Result: not yet measured&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;Directional correlation only — does an attack that fires in sim also fire on hardware, and does the benign control stay near zero on both. With n = 5 the real intervals are wide by design; no real-robot ASR is claimed, before or after.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/sim-to-real/&quot;&gt;The protocol&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>pre-registered</category></item><item><title>MEASURED: Instruction canonicalization lowered ASR without breaking the task</title><link>https://www.provael.com/findings#defense-mitigation</link><guid isPermaLink="true">https://www.provael.com/findings#defense-mitigation</guid><description>adversarial ASR 67.5% [52-80%] -&gt; 7.5% [3-20%] on stub · instruction 70.0% [52-83%] -&gt; 10.0% [3-26%] · optimized_instruction 60.0% [31-83%] -&gt; 0.0% [0-28%] · benign FPR 0.0% -&gt; 0.0% · clean-task success 100.0% -&gt; 100.0% (within CI, accepted)</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;defense · mitigation&lt;/strong&gt; — MEASURED&lt;/p&gt;&lt;p&gt;adversarial ASR 67.5% [52-80%] -&amp;gt; 7.5% [3-20%] on stub · instruction 70.0% [52-83%] -&amp;gt; 10.0% [3-26%] · optimized_instruction 60.0% [31-83%] -&amp;gt; 0.0% [0-28%] · benign FPR 0.0% -&amp;gt; 0.0% · clean-task success 100.0% -&amp;gt; 100.0% (within CI, accepted)&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Defense: instruction canonicalization / repair&lt;/li&gt;&lt;li&gt;Target: EAI01 instruction · EAI01 optimized&lt;/li&gt;&lt;li&gt;Suites: stub + reach (CPU fixture)&lt;/li&gt;&lt;li&gt;Gate: accepted&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;&lt;em&gt;Credited, and substantially circular - the study says so above its own results table. Four of the CPU fixture’s seven danger tokens are words this defense strips, so the optimized_instruction result is close to tautological. stub-validated scaffolding; no real-model transfer is claimed, and the reach arm’s acceptance gate was not evaluable.&lt;/em&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://github.com/provael/provael/blob/main/docs/studies/instruction-canonicalization.md&quot;&gt;The study ↗&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.provael.com/eai-top-10/eai01/&quot;&gt;How EAI01 is tested&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content:encoded><category>findings</category><category>measured</category></item></channel></rss>