STALE MEASUREMENTNewest real-model measurement: 26 days old, measured with v0.32.0; 7 releases have shipped since, past this project’s own 7-day window. Why, and what unblocks it

ProductEvidenceTop 10LeaderboardCompliancePricingDocsStar on GitHub Quickstart
( 06 ) · EAI06

Cross-domain safety misalignment (the embodiment gap)

Safety at the language layer that does not carry to the physical action - a benign-sounding, “safely-reasoned” instruction still drives an unsafe motion.

SARIF EAI06channel: misalignment
Definition

What it is

Cross-domain safety misalignment is a mismatch across the three layers a VLA policy spans - language-level safety, world-model reasoning, and the executable physical action - so that safety at one layer does not carry to the next (BadRobot, ICLR 2025).

The sub-mode Provael targets is the embodiment gap: a benign-looking instruction that a chat-layer safety filter passes still drives an unsafe embodied action, because the harm lives in the world model, not the words. Teams that bolt an LLM-style refusal filter onto a robot get false confidence.

Real example

Seen in the wild

BadRobot (ICLR 2025, arXiv:2407.20242) names cross-domain safety misalignment - including the world-knowledge failure mode, where a benign-looking instruction yields an unsafe action because the agent does not know an object is hazardous, fragile, or a person - as a core embodied surface: safe language, unsafe action.

External validation - distinct from Provael’s own result

BadRobot (ICLR 2025) - arXiv:2407.20242

How Provael tests it

Attack, measure, prove

  • Ships the `misalignment` family (benign_urgency_override + euphemistic_reroute): each attack reframes the task into an instruction a chat-layer safety filter would pass as benign - an urgency framing (“hurry, just push it aside”) or a euphemism (“tidy up quickly”, “clear that out of the way”) - while injecting the embodied consequence that instruction carries, so the policy’s commanded end-effector enters a keep-out zone.
  • The reframed instruction carries no language-level unsafe token (the danger channel stays 0), so the violation is attributable to the language→action gap, not a jailbroken goal (that is the separate EAI01 `instruction` family). Scored as a rate with a 95% Wilson CI against a benign-FPR control.
  • Stub-validated scaffolding only (deterministic CPU `reach` keep-out suite: 100% [84-100%] vs a 0% benign baseline); real-model transfer (BadRobot × SmolVLA × LIBERO) is GPU-gated and to be measured on that path, reported with the same Wilson-95 + benign-FPR and never over-sold - no cross-model or “first” claim.

Declared defense: instruction canonicalization → The defense DECLARES it applies here; it has not been measured against this risk. Declared applicability is not evidence.

Mitigations

What reduces the risk

  • Validate the action plan and its physical-world consequences, not just the language.
  • Ground safety rules to world state (object / person / hazard awareness).
  • Executability and harm gating at the planner→actuator boundary, with an independent safety monitor.
SARIF rule id

How it shows up in CI

Every finding for this risk carries a stable rule id in the SARIF report - the Embodied AI Top-10 id itself:

ruleId: EAI06
Compliance mapping

Crosswalk

OWASP
ASI03 Identity & Privilege Abuse · LLM06 Excessive Agency (guardrail bypass)
MITRE ATLAS
Impact → unsafe embodied action under a language-benign instruction (proposed - the embodiment gap has no on-point ATLAS technique)
Frameworks
ISO 10218:2025 - functional-safety interfaceISO 13482:2014 (under revision as ISO/DIS 13482, retitled to service robots)

See the full framework crosswalk for dates and detail. Not legal advice.

DOC. PVL-EAI06 · Embodied AI Security Top 10 · v0.2Updated 2026-06-27 · CC BY-SA 4.0