How the risk scores work

Conformal duration-anomaly scoring, real policy failure rates, and how we verified both before shipping them.

What it actually does

Every episode's Duration anomaly card uses split conformal prediction to score how unusual that episode's duration is compared to successful runs of the same task, with a real distribution-free coverage guarantee.

Concretely: it computes a p-value answering "if this episode came from the same distribution as our successful calibration runs, how surprising would a duration this extreme be?" A low p-value means the duration looks statistically unlike the successful runs we've seen: worth a human's attention, not proof of anything on its own.

The same scoring rolls up per policy version too: instead of asking whether one past recording looked odd, it asks how often a given deployed policy's logged runs have looked unusual so far, compared to other runs of the exact same task. It is still a rollup of the same after-the-fact duration scores, not a new kind of prediction.

Policy failure rate

The Policy failure ratetable on the Overview page is a different kind of number: a named policy checkpoint's own real, measured pass/fail record, pooled across every task it has been independently evaluated on, with a Wilson confidence interval (Wilson, 1927) on the true rate. It comes from RoboArena's public evaluation dump (MIT-licensed, arXiv:2506.18123): real autonomous rollouts of named VLA policy checkpoints on a real DROID/Franka robot, each scored pass or fail by a human evaluator, not from our own fleet.

This is the one number on this platform that is genuinely forward-looking, and it is forward-looking only in a narrow, specific sense: an unmodified policy checkpoint's behavior on similar tasks is reasonably estimated by its own prior track record, since the same weights produce the same kind of behavior run to run. It describes that exact checkpoint, evaluated by a third party, on their tasks and their robot. It is not a prediction about any specific customer's hardware, environment, or deployment, and a checkpoint with a good measured rate elsewhere can still fail differently on a materially different setup.

What it does not do

Neither the per-episode duration score nor the per-task policy rollup predicts robot failure before it happens. Both are purely retrospective: they describe how past logged runs looked, not how a policy will perform on a run that hasn't happened yet. Real intervention/collision telemetry is not populated on any dataset we have, which rules out several other honest signals we'd otherwise build.

The policy failure rate above is the exception, and only in the narrow sense described there: it is still not a claim about a specific customer's deployment, and a low failure rate on RoboArena's benchmark tasks is not a guarantee on tasks that checkpoint has never been evaluated on.

How we tested it

Automated tests back every piece of this: unit tests on the underlying math (conformal p-values, the nonconformity scoring formula, the Wilson interval against a standard worked example, edge cases like empty or degenerate reference sets), and integration tests that run against the real dataset, not synthetic data.

  • Coverage guarantee test: confirms that on real held-out successful episodes, scored through the same per-task calibration the product actually uses, the empirical false-positive rate stays near the nominal alpha threshold.
  • Discriminative-power test: confirms that real labeled failures score as more anomalous, on average, than held-out real successes. This is a falsifiable claim: if episode duration carried zero signal about failure, this test would fail. It passes.
  • Wilson interval tests: check the formula against a standard published worked example (5 successes of 10 trials), confirm the interval always contains the observed rate and stays within [0, 1], and confirm it narrows as sample size grows.

Methodology provenance

Duration-anomaly scoring is built on "A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification" (Angelopoulos & Bates, arXiv:2107.07511). Before building on it, 30 candidate ML papers were researched and adversarially re-verified: each arXiv link and any claimed code license was checked directly rather than trusted from the initial research pass. This paper was selected because its MIT license was confirmed directly against the repository, and the underlying algorithm is simple enough to implement independently either way.

Policy failure rate data comes from RoboArena (arXiv:2506.18123), whose public dataset dump is MIT-licensed, verified directly against the dataset's own license metadata before use. The confidence interval uses the Wilson score interval (Wilson, E.B., 1927, "Probable Inference, the Law of Succession, and Statistical Inference"), the standard correction for the normal approximation interval's known failure at small sample sizes or near 0% and 100%.