# EAI10: Insufficient evaluation, observability & incident response

> The governance meta-risk: no standardized eval, no telemetry to detect an attack, no rollback, no incident process — so the other nine go unmanaged.

Out of scope for a VLA-policy red-teamer by design — no Provael attack, no SARIF rule.

## Definition

The meta-risk: no standardized real-world safety eval, no logging/telemetry to detect an attack, no policy rollback, and no incident process — so the other nine risks go undetected and unmanaged. You cannot defend, detect, or recover from what you cannot measure or see.

It is out of scope for Provael as an attack target — it is a governance meta-risk that Provael’s own evaluation mitigates rather than attacks. Running a pre-deploy red-team with an ASR scorecard (n, CIs, clean baseline) is part of the answer to EAI10, not a thing to attack.

## Real example

a16z’s “The Physical AI Deployment Gap” puts it bluntly: the robotics equivalent of DevOps practices does not exist yet — missing observability, failure-mode testing, and behavioral characterization. No agreed embodied-AI incident-reporting framework exists yet.

## How Provael tests it

- Provael does not attack EAI10 — it is a governance and observability meta-risk, not a VLA-policy behaviour. There is no attack family, no channel, and no SARIF rule for it.
- Instead, Provael is part of the mitigation: its pre-deploy red-team produces exactly the ASR scorecard (with n, confidence intervals, and a clean benign baseline) that EAI10 says is missing.
- The rest is operational: runtime observability on the perception→action path, OOD / distribution-shift monitors, versioned policies with rollback, and a written incident plan.

## Mitigations

- Pre-deploy red-team and an ASR scorecard with n, CIs, and a clean baseline.
- Runtime observability on the perception→action path; OOD / distribution-shift monitors; graceful hand-back-to-human.
- Versioned policies with rollback, and a written disclosure + incident plan.

## Crosswalk

- OWASP: OWASP ASI — inadequate monitoring (no numbered peer)
- MITRE ATLAS: No ATLAS technique — governance meta-risk
- Frameworks: NIST AI RMF — Measure / Manage, ISO 10218:2025 — validation

Canonical: https://www.provael.com/eai-top-10/eai10 · v0.2

---
Provael · Prove it. Prevail. · Apache-2.0 · https://github.com/provael/provael
Not legal advice; verify regulatory dates against the primary source.
