Repository navigation
E56: on Hopper, the step ledger catches and locates every fault in what crosses; curve checks miss all 69 at one value - #123
Merged
Conversation
…opper, which ends episodes when it falls, do E54's and E55's findings hold, and what do faults in terminations do E54's 21 client faults plus terminated:drop and terminated:as-trunc, and the 2 log faults, at doses one, 0.01 and 1.0; seeds 0-2 traced and ledgered, seeds 10-19 fault-free; 244 runs. Status rule: last tenth of the log beats the first tenth by 1,000.
…p ledger catches and locates all 207 client-fault runs, and the curve checks miss every one at one value Registered runs, 244 of 244 complete; every check and prediction holds. Status rule 0/3/33 of 69 client-fault runs at doses one/0.01/1.0; byte identity 57/69/69 (it misses the four final-obs faults at one value, which change a final observation PPO never reads); ledger 69/69/69 with no false alarm in 27 control and log-fault runs; ledger or record 225 of 225. verdicts.py is committed unchanged; FINDINGS.md says when it was written.
…ated on a clean environment, is a policy trained through a boundary fault worse, and are the faults both curve checks miss? Every E54 and E56 final policy (834) plays 50 deterministic episodes in a fresh gymnasium env, seeds 10000-10049. Harmed = clean return more than 3 s.d. below the fault-free band (seeds 10-19). Predictions: few false alarms (P1), no measurable harm at one value (P2), Pendulum reward:zero harmed behind a perfect-looking log (P3). evaluate.py, run.sh, summarise.py and verdicts.py are committed here, before any registered run.
…s that both curve checks miss are as good as fault-free ones; every policy the faults wrecked, the curve checks had already caught 834 of 834 evaluations; V1, V2, P1-P3 hold. Harmed while silent on the curve: 5 of 255 runs on Pendulum (below all 13 fault-free runs), 0 on HalfCheetah and Hopper. Exploratory: silent runs sit 0.03-0.65 band s.d. above the band. Pendulum reward:zero logs 0 (best possible) while its policies score -1196 to -1639. One amendment: a parsing fix in summarise.py for the loud faults' blank cells, made before any output.
E57: on a clean environment, faults the curve checks miss leave policies as good as fault-free ones; every wrecked policy was already caught
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The question
E54 and E55 injected faults on Pendulum and HalfCheetah, which never end an episode early. Hopper ends an episode when it falls, and it is harder to learn. On Hopper, do E54's and E55's findings hold? And what do faults in terminations do?
The answer
They hold. The faults were E54's 21 client faults plus two in terminations (
terminated:drop,terminated:as-trunc) and the two log faults, at doses one, 0.01 and 1.0. There were 244 runs of 1,024,000 steps.Client-fault runs with a changed value, 69 at each dose:
Every check and prediction holds (
results/verdicts.txt).final-obsfault changed the final observation of an episode that ended in a fall. SB3 bootstraps from it only at a time limit, so the weights stayednone's in 12 of 12 runs. Only the ledger caught them. At 0.01 and 1.0 the changed values land at time limits too, and byte identity caught 24 of 24.terminated:as-truncchanges nothing in the record's books, so the record missed it at every dose. Withterminated:dropat 1.0, the trainer's log held no episode, while the environment's record held 7,178.verdicts.pywas not in the protocol commit. It was last changed after the registered runs started and before the first one ended. FINDINGS.md gives the times, and the file is committed unchanged.The per-run directories stay on guangzhao. Their traces and ledger arrays (150 MB a run) were deleted once judged, as
run.shsays.