Skip to content

Repository files navigation

extend-SAFE

Pooled ROC-AUC (0.747 / 0.755) does not imply reliable early failure warning on unseen tasks. Under 10-fold LOTO with SAFE functional conformal calibration at ≤5% realized FPR, catch rate is only 8.5% (MLP) / 6.9% (LSTM) with effective mean alarm positions of 95.5% / 95.7% of the evaluation window (missed failures count as alarm-at-end).

Failure catch rate versus realized false-positive rate. Color records effective alarm position — lower (greener) is earlier.

This repository is a reproducible audit of SAFE-style failure detection over OpenVLA representations on ten LIBERO tasks. Its main finding is deliberately narrow: in this reproduction, strong pooled ROC-AUC does not imply timely, low-false-positive warning on an unseen task.

On the 1,000-rollout corpus, the three original task-held-out splits obtain mean pooled early-window AUCs of 0.747 (MLP) and 0.755 (LSTM). Separately, ten-fold equal-task LOTO AUCs are 0.672 and 0.652. At the best tested SAFE functional-conformal operating points with at most 5% realized FPR, they catch only 8.5% and 6.9% of held-out-task failures.

Read the hosted audit, the GitHub-rendered technical report, its Quarto source, or the explicit closeout requirements.

Reproduce the public audit

Install the exact tested environment:

uv sync --frozen --extra dev

The included public score bundle contains detector trajectories and rollout metadata. It replays the original and fixed-horizon functional-CP analyses without raw hidden states or checkpoints:

uv run python scripts/evaluate_functional_cp_loto.py \
  --score-bundle-in docs/results_audit/score_bundle \
  --output-dir reproduced/functional_cp_loto \
  --fixed-horizons 50,100,148

uv run python scripts/compare_functional_cp_outputs.py \
  --expected-root docs/results_audit \
  --actual-root reproduced/functional_cp_loto

Validate code and publication artifacts with:

uv run pytest -q
uv run python scripts/verify_publication.py

See DATA.md for the full-inference input contract and the score-bundle specification for the compact replay format.

Scope

  • Primary result: 1,000 rollouts, 100 per task, ten LOTO folds.
  • Policy/benchmark: OpenVLA on ten LIBERO tasks.
  • Benchmark suite: libero_10, confirmed from saved rollout metadata.
  • Primary monitors: SAFE-style layer-32 MLP and LSTM probes.
  • Secondary layer-20/32 result: 1,000 rollouts, 100 per task, ten LOTO folds, and three training seeds. Capacity-matched fusion does not improve AUC (-0.004; 95% task-bootstrap CI [-0.020, 0.012]) or fixed-horizon warning. A residual model improves one retrospective warning operating point, but not the predeclared fixed horizons, so it remains exploratory. The original 450-rollout, nine-task layer-24/28/32 signal now has an independent 500-rollout, ten-task replication: its AUC point estimates reproduce, but the direct interval remains unresolved and warning utility does not improve.
  • The repository does not claim that SAFE fails for every policy, benchmark, or failure-label regime.

Project status

The main scientific conclusion is stable at the scope above, and open limitations and non-goals are recorded in the report. The public score bundle replays the original and fixed-horizon functional-CP analyses; the predeclared 50/100/148-step sensitivity is included. Compact result tables live under docs/results_audit/; raw rollouts, checkpoints, and generated development runs are excluded. Two small historical records remain under runs/openvla_layers/, explicitly separate from the final audit's evidence. The fixed-horizon results retain the negative warning conclusion: at tested points meeting the 5% macro FPR cap, MLP catches 3.6%, 3.5%, and 6.8% at 50/100/148 steps. LSTM has no qualifying tested point at 50 or 100 steps and catches 7.0% at 148 steps.

The rollouts were independently collected by the maintainer using the same OpenVLA weights and seeds as SAFE's authors, with raw data backed up on Google Drive. This release supports score-level replay; the original collection command and immutable policy-weight revision are not archived here. The upstream collection recipe, evidence for task identities, recovered probe settings, and training commands are in EXPERIMENT_PROVENANCE.md.

Edit docs/safe_openvla_audit.qmd, then render with Quarto 1.9.37 and rebuild the publication manifest after all other edits:

quarto render docs/safe_openvla_audit.qmd --to gfm
uv run python scripts/build_publication_manifest.py
uv run python scripts/verify_publication.py

CI renders the Quarto source and checks that the tracked Markdown matches. On main, it also builds and deploys the GitHub Pages site after all tests, artifact checks, and score replay pass. Build the same site locally with uv run python scripts/build_site.py; its output is reproduced/pages/_site/. The builder checks every local page, asset, download, and HTML anchor before deployment.

License and citation

Original repository code is released under the MIT License. Third-party models, benchmarks, papers, and data retain their own terms; see THIRD_PARTY.md. Citation metadata are in CITATION.cff.

About

Reproducible audit of SAFE VLA failure detection on OpenVLA/LIBERO — pooled AUC inflates +0.09 over macro within-task; functional CP catches only 7-9% at ≤5% FPR

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages