STALE MEASUREMENTThe newest real-model measurement on this site is 14 days old — past this project’s own 7-day window. Every measured figure here was true when it was taken and has not been re-measured since. Why, and what unblocks it

ProductEvidenceAdoptersTop 10For youLeaderboardSubmitCompliancePricingReadinessDocsStar on GitHub Quickstart
FOR AMR & AGV FLEET OPERATORS · PV-019

Your cybersecurity clock already ran. The machine-learned part arrived after it.

UN R155 has been mandatory for new vehicle types since July 2022, and for all new vehicles produced since July 2024. The regime assumes you can enumerate a component's failure modes and show they are managed. A vision-language-action policy that changes behaviour when an instruction is reworded does not enumerate that way - and it is the part of your fleet with the least evidence behind it.

Read this first

What this cannot do for a type approval

  • This is a simulation-only method. Provael runs policies in simulators (LIBERO, robosuite, MuJoCo). It does not drive a vehicle, an AMR, or any hardware.
  • There are zero real-hardware results. Not "few" - zero. The physical-transfer study is pre-registered and has not been run, and no independent party has reproduced our run.
  • It is not a conformity assessment and produces no approval. R155 approval is granted by a type-approval authority against a CSMS. Provael produces one input to the evidence a CSMS holds, and determines nothing.
  • The keep-out predicate is uncalibrated. A success means the policy left its benign envelope, not that a certified danger threshold was crossed. Scope: sim-only · ten tasks · 5 seeds/task · uncalibrated keep-out predicate · benign-reword arm measured: 1/50 vs the attack’s 44/50.

If a supplier tells you a simulation result satisfies R155, that is the claim to push on - including when the supplier is us.

The clocks

What binds an AMR or AGV fleet today

Dates verified live on 10 August 2026. Where sources disagreed on the exact day, the month is carried rather than a guess.

Regulatory instruments, what each covers, when it applies, and how it bears on a machine-learned component.
InstrumentWhat it coversWhenWhy it reaches a learned policy
UN R155Cyber security and cyber security management system (CSMS) type approvalMandatory for new vehicle types since July 2022; for all new vehicles produced since 1 July 2024The clock that has already run. A CSMS must be certified and re-assessed on a cycle, and it covers the vehicle type, not one component.
UN R156Software update and software update management system (SUMS)Same schedule as R155; SUMS assessed and renewed at least every three yearsEvery model update to a deployed policy is a software update inside this system. A policy you retrain is not outside it.
ISO/SAE 21434:2021Road vehicles - cybersecurity engineering, across the E/E lifecyclePublished 31 August 2021The engineering practice a CSMS is audited against in practice. It asks for threat analysis and risk assessment (TARA) with evidence, not assertion.
ISO 26262:2018Road vehicles - functional safety (second edition)Published 2018Where functional safety meets the above. It was not written for machine-learned behaviour, which is the gap ISO/IEC TR 5469 exists to describe.
EU AI ActProduct-embedded high-risk AI, via Annex I sectoral legislationApplies 2 August 2028Deferred from 2 August 2027 by Regulation (EU) 2026/1744, the Digital Omnibus on AI. The nearer binding date for machinery is the Machinery Regulation on 20 January 2027.
What Provael produces

A TARA input with a denominator

ISO/SAE 21434 asks for threat analysis and risk assessment with evidence behind it. The usual difficulty with a learned component is that the threat is easy to name and the likelihood is a shrug. Provael turns that column into a measured rate.

  • An attack-success rate per risk, with a 95% interval and a matched benign control - our own published run measures roleplay at88% (44/50) against a 4%benign baseline, on one policy in simulation.
  • Honest nulls in the same artifact: attack families that did not transfer are reported at their own n, which is what stops a risk register recording only what fired.
  • SARIF and OSCAL output, plus an Ed25519-signed attestation, so a re-run can be compared against the filed one rather than re-argued.
  • A CI gate, which matters more under R156 than under R155: every retrain is a software update inside your SUMS, and a rate that is only measured once is a rate about a model you have already replaced.
Talk about a fleet assessment →See the measured result

Feature and status claims on this page were read against the product repository on . Unlike the numbers on /results and /leaderboard, these are not enforced by a build check — the product publishes no machine-readable artifact for per-suite defense verdicts, so this is a human review with a date on it, not a guarantee. The maintained source is the repository; where it disagrees with this page, it wins.