Pooled ROC-AUC (0.747 / 0.755) does not imply reliable early failure warning on unseen tasks. Under 10-fold LOTO with SAFE functional conformal calibration at ≤5% realized FPR, catch rate is only 8.5% (MLP) / 6.9% (LSTM) with effective mean alarm positions of 95.5% / 95.7% of the evaluation window (missed failures count as alarm-at-end).
This repository is a reproducible audit of SAFE-style failure detection over OpenVLA representations on ten LIBERO tasks. Its main finding is deliberately narrow: in this reproduction, strong pooled ROC-AUC does not imply timely, low-false-positive warning on an unseen task.
On the 1,000-rollout corpus, the three original task-held-out splits obtain mean pooled early-window AUCs of 0.747 (MLP) and 0.755 (LSTM). Separately, ten-fold equal-task LOTO AUCs are 0.672 and 0.652. At the best tested SAFE functional-conformal operating points with at most 5% realized FPR, they catch only 8.5% and 6.9% of held-out-task failures.
Read the hosted audit, the GitHub-rendered technical report, its Quarto source, or the explicit closeout requirements.
Install the exact tested environment:
uv sync --frozen --extra devThe included public score bundle contains detector trajectories and rollout metadata. It replays the original and fixed-horizon functional-CP analyses without raw hidden states or checkpoints:
uv run python scripts/evaluate_functional_cp_loto.py \
--score-bundle-in docs/results_audit/score_bundle \
--output-dir reproduced/functional_cp_loto \
--fixed-horizons 50,100,148
uv run python scripts/compare_functional_cp_outputs.py \
--expected-root docs/results_audit \
--actual-root reproduced/functional_cp_lotoValidate code and publication artifacts with:
uv run pytest -q
uv run python scripts/verify_publication.pySee DATA.md for the full-inference input contract and the score-bundle specification for the compact replay format.
- Primary result: 1,000 rollouts, 100 per task, ten LOTO folds.
- Policy/benchmark: OpenVLA on ten LIBERO tasks.
- Benchmark suite:
libero_10, confirmed from saved rollout metadata. - Primary monitors: SAFE-style layer-32 MLP and LSTM probes.
- Secondary layer-20/32 result: 1,000 rollouts, 100 per task, ten LOTO folds, and three training seeds. Capacity-matched fusion does not improve AUC (-0.004; 95% task-bootstrap CI [-0.020, 0.012]) or fixed-horizon warning. A residual model improves one retrospective warning operating point, but not the predeclared fixed horizons, so it remains exploratory. The original 450-rollout, nine-task layer-24/28/32 signal now has an independent 500-rollout, ten-task replication: its AUC point estimates reproduce, but the direct interval remains unresolved and warning utility does not improve.
- The repository does not claim that SAFE fails for every policy, benchmark, or failure-label regime.
The main scientific conclusion is stable at the scope above, and open
limitations and non-goals are recorded in the report. The public score bundle
replays the original and fixed-horizon functional-CP analyses; the
predeclared 50/100/148-step sensitivity is included. Compact result tables live
under docs/results_audit/; raw rollouts, checkpoints, and generated development
runs are excluded. Two small historical records remain under
runs/openvla_layers/, explicitly separate from the final audit's evidence.
The fixed-horizon results retain the negative warning conclusion: at tested
points meeting the 5% macro FPR cap, MLP catches 3.6%, 3.5%, and 6.8% at
50/100/148 steps. LSTM has no qualifying tested point at 50 or 100 steps and
catches 7.0% at 148 steps.
The rollouts were independently collected by the maintainer using the same OpenVLA weights and seeds as SAFE's authors, with raw data backed up on Google Drive. This release supports score-level replay; the original collection command and immutable policy-weight revision are not archived here. The upstream collection recipe, evidence for task identities, recovered probe settings, and training commands are in EXPERIMENT_PROVENANCE.md.
Edit docs/safe_openvla_audit.qmd, then render with Quarto 1.9.37 and rebuild
the publication manifest after all other edits:
quarto render docs/safe_openvla_audit.qmd --to gfm
uv run python scripts/build_publication_manifest.py
uv run python scripts/verify_publication.pyCI renders the Quarto source and checks that the tracked Markdown matches.
On main, it also builds and deploys the GitHub Pages site after all tests,
artifact checks, and score replay pass. Build the same site locally with
uv run python scripts/build_site.py; its output is
reproduced/pages/_site/. The builder checks every local page, asset, download,
and HTML anchor before deployment.
Original repository code is released under the MIT License. Third-party models, benchmarks, papers, and data retain their own terms; see THIRD_PARTY.md. Citation metadata are in CITATION.cff.
