Physical AI · evidence and evaluation

Runtime monitoring, intervention and what the system does when it is wrong

Behaviour on failure is a design choice that should be specified before deployment: what is monitored, what threshold triggers, what the system does next, and who can intervene.

How it works

NIST publishes its AI Risk Management Framework for voluntary use across design, development, use and evaluation, which is the right altitude for this: it prompts the questions without deciding them. A monitor is only meaningful together with a fallback — stop, hold position, hand back to an operator, retry with reduced speed — and with an intervention path that someone can actually reach in time. Maha’s proposal is to write the monitor, the threshold, the fallback and the intervention route into the evaluation contract, and then to report how often each fired, including false alarms, because a monitor nobody trusts gets switched off.

A concrete case

A force threshold that halts on contact will also halt on a normal insertion. If the recorded evaluation does not separate “halted on a real anomaly” from “halted on an ordinary contact”, the operator’s experience of nuisance stops is invisible in the result.

What this establishes

That a recognised public framework exists for organising these questions. Nothing about any specific monitor’s performance.

What it does not

Following a voluntary framework certifies nothing and supplies no acceptance criterion. Safety engineering for a physical machine requires qualified people and applicable standards.

Questions worth asking

  • Ask what is monitored, at what rate, and what the fallback state is.
  • Ask for the counts of true and false triggers during evaluation.
  • Ask how an operator intervenes, how long that takes, and what the system does meanwhile.

Sources

  • NIST AI Risk Management Framework

    NIST states that the AI Risk Management Framework is intended for voluntary use and to improve the ability to incorporate trustworthiness considerations into the design, development, use and evaluation of AI products, services and systems.

    Boundary: Voluntary guidance, and following it certifies nothing. The framework is general-purpose rather than prescriptive for a domain: NIST has issued companion profiles for other sectors, but there is no robot-specific profile and no acceptance criterion for a physical machine anywhere in it.

    Locator, anchor and reuse basis

    Read at: Overview of the AI RMF — the framework’s own purpose statement. Inspected 2026-09-20. Original paraphrase and link only.

    Verify by searching the source for: intended for voluntary use. If that phrase is not there, or does not carry the meaning stated above, this citation is wrong and we want to know.

Continue

Elsewhere on this site

Maha Strategies publishes explanation and evaluation method. We build no robots, run no physical experiments, and report no benchmark results of our own. Hardware, safety and evidence-intake questions live in the robotics section.