Skip to content

E56: on Hopper, the step ledger catches and locates every fault in what crosses; curve checks miss all 69 at one value - #123

Merged
tactino merged 5 commits into
mainfrom
exp/e56-hopper-faults
Oct 3, 2026
Merged

tactino merged 5 commits into
mainfrom
exp/e56-hopper-faults

Conversation

@tactino

@tactino tactino commented Oct 3, 2026

Copy link
Copy Markdown
Member

The question

E54 and E55 injected faults on Pendulum and HalfCheetah, which never end an episode early. Hopper ends an episode when it falls, and it is harder to learn. On Hopper, do E54's and E55's findings hold? And what do faults in terminations do?

The answer

They hold. The faults were E54's 21 client faults plus two in terminations (terminated:drop, terminated:as-trunc) and the two log faults, at doses one, 0.01 and 1.0. There were 244 runs of 1,024,000 steps.

Client-fault runs with a changed value, 69 at each dose:

dose status rule strict byte identity record step ledger
one value 0 0 57 19 69
0.01 3 3 69 21 69
1.0 33 34 69 18 69

Every check and prediction holds (results/verdicts.txt).

  • One changed value moved the final return by a median of 197, against a fault-free seed s.d. of 259. Neither curve check caught any of the 69 runs.
  • At 1.0, 11 faults that change every value still passed both curve checks at all three seeds.
  • The ledger caught and located all 207 client-fault runs. It found no difference in the 27 control and log-fault runs. The strict check flagged none of the 10 fault-free seeds.
  • New on Hopper: a fault can cross without changing the weights. At one value, each final-obs fault changed the final observation of an episode that ended in a fall. SB3 bootstraps from it only at a time limit, so the weights stayed none's in 12 of 12 runs. Only the ledger caught them. At 0.01 and 1.0 the changed values land at time limits too, and byte identity caught 24 of 24.
  • terminated:as-trunc changes nothing in the record's books, so the record missed it at every dose. With terminated:drop at 1.0, the trainer's log held no episode, while the environment's record held 7,178.

verdicts.py was not in the protocol commit. It was last changed after the registered runs started and before the first one ended. FINDINGS.md gives the times, and the file is committed unchanged.

The per-run directories stay on guangzhao. Their traces and ledger arrays (150 MB a run) were deleted once judged, as run.sh says.

…opper, which ends episodes when it falls, do E54's and E55's findings hold, and what do faults in terminations do

E54's 21 client faults plus terminated:drop and terminated:as-trunc, and the 2 log faults, at doses one, 0.01 and 1.0; seeds 0-2 traced and ledgered, seeds 10-19 fault-free; 244 runs. Status rule: last tenth of the log beats the first tenth by 1,000.
…p ledger catches and locates all 207 client-fault runs, and the curve checks miss every one at one value

Registered runs, 244 of 244 complete; every check and prediction holds. Status rule 0/3/33 of 69 client-fault runs at doses one/0.01/1.0; byte identity 57/69/69 (it misses the four final-obs faults at one value, which change a final observation PPO never reads); ledger 69/69/69 with no false alarm in 27 control and log-fault runs; ledger or record 225 of 225. verdicts.py is committed unchanged; FINDINGS.md says when it was written.
…ated on a clean environment, is a policy trained through a boundary fault worse, and are the faults both curve checks miss?

Every E54 and E56 final policy (834) plays 50 deterministic episodes in a fresh gymnasium env, seeds 10000-10049. Harmed = clean return more than 3 s.d. below the fault-free band (seeds 10-19). Predictions: few false alarms (P1), no measurable harm at one value (P2), Pendulum reward:zero harmed behind a perfect-looking log (P3). evaluate.py, run.sh, summarise.py and verdicts.py are committed here, before any registered run.
…s that both curve checks miss are as good as fault-free ones; every policy the faults wrecked, the curve checks had already caught

834 of 834 evaluations; V1, V2, P1-P3 hold. Harmed while silent on the curve: 5 of 255 runs on Pendulum (below all 13 fault-free runs), 0 on HalfCheetah and Hopper. Exploratory: silent runs sit 0.03-0.65 band s.d. above the band. Pendulum reward:zero logs 0 (best possible) while its policies score -1196 to -1639. One amendment: a parsing fix in summarise.py for the loud faults' blank cells, made before any output.
E57: on a clean environment, faults the curve checks miss leave policies as good as fault-free ones; every wrecked policy was already caught
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant