ProductEvidenceTop 10ComplianceDocsStar on GitHubQuickstart
( 10 ) · EAI10

Insufficient evaluation, observability & incident response

The governance meta-risk: no standardized eval, no telemetry to detect an attack, no rollback, no incident process — so the other nine go unmanaged.

out of scope for a VLA-policy red-teamer
Definition

What it is

The meta-risk: no standardized real-world safety eval, no logging/telemetry to detect an attack, no policy rollback, and no incident process — so the other nine risks go undetected and unmanaged. You cannot defend, detect, or recover from what you cannot measure or see.

It is out of scope for Provael as an attack target — it is a governance meta-risk that Provael’s own evaluation mitigates rather than attacks. Running a pre-deploy red-team with an ASR scorecard (n, CIs, clean baseline) is part of the answer to EAI10, not a thing to attack.

Real example

Seen in the wild

a16z’s “The Physical AI Deployment Gap” puts it bluntly: the robotics equivalent of DevOps practices does not exist yet — missing observability, failure-mode testing, and behavioral characterization. No agreed embodied-AI incident-reporting framework exists yet.

Scope

Why this is out of scope

  • Provael does not attack EAI10 — it is a governance and observability meta-risk, not a VLA-policy behaviour. There is no attack family, no channel, and no SARIF rule for it.
  • Instead, Provael is part of the mitigation: its pre-deploy red-team produces exactly the ASR scorecard (with n, confidence intervals, and a clean benign baseline) that EAI10 says is missing.
  • The rest is operational: runtime observability on the perception→action path, OOD / distribution-shift monitors, versioned policies with rollback, and a written incident plan.
Mitigations

What reduces the risk

  • Pre-deploy red-team and an ASR scorecard with n, CIs, and a clean baseline.
  • Runtime observability on the perception→action path; OOD / distribution-shift monitors; graceful hand-back-to-human.
  • Versioned policies with rollback, and a written disclosure + incident plan.
Compliance mapping

Crosswalk

OWASP
OWASP ASI — inadequate monitoring (no numbered peer)
MITRE ATLAS
No ATLAS technique — governance meta-risk
Frameworks
NIST AI RMF — Measure / ManageISO 10218:2025 — validation

See the full framework crosswalk for dates and detail. Not legal advice.

DOC. PVL-EAI10 · Embodied AI Security Top 10 · v0.2Updated 2026-06-27 · CC BY-SA 4.0