Skip to content

E54: a fault at the boundary passes the status rule almost always; byte-identical weights or an env-side record catches every one - #121

Merged
tactino merged 2 commits into
mainfrom
exp/e54-boundary-faults
Oct 2, 2026
Merged

tactino merged 2 commits into
mainfrom
exp/e54-boundary-faults

Conversation

@tactino

@tactino tactino commented Oct 2, 2026

Copy link
Copy Markdown
Member

The question

Across machines, E43, E44, E48 and E49 judge a boundary by a status rule: the policy still learns. When something goes wrong at the environment-trainer boundary, does that rule notice? What does?

The answer

Almost never: the status rule misses nearly every fault. Byte-identical weights, or reconciliation with a record kept at the environment, catch every one.

SB3's PPO trained Pendulum and HalfCheetah through the bridge, as in E50. Meanwhile the env client changed one field of what crossed, in one way: 21 faults, covering every field with zero, stale, swap, float16 and clip, plus episode-end and step faults. Two more faults sat in the trainer's log. Doses ran from one value to every value. 590 registered runs, seeds 0-2.

env dose status rule strict (10 seeds) weights record weights or record
Pendulum one value 0% 0% 91% 28% 100%
1% 4% 6% 91% 35% 100%
every value 39% 55% 91% 30% 100%
HalfCheetah one value 4% 4% 91% 33% 100%
every value 32% 46% 91% 33% 100%

Every check and prediction holds (results/verdicts.txt).

  • Weights caught all 504 client-fault runs at every dose, and no control or log fault. The 91% in the table is client and log faults together.
  • Record caught what changes the books: rewards beyond its tolerance, episode ends, steps, the log. It did not catch faults in observations or actions.
  • Between them, the two checks caught all 546 silent runs.
  • Localisation: the first difference from the fault-free run named the step, env and field in all 351 traced runs.
  • Loud faults: the two faults the bridge rejects raised at once.

One changed value moved the final return by about as much as a change of seed. That is why a curve cannot tell a fault from a seed.

The per-run logs (100 MB) stay on guangzhao. results/runs.jsonl and runs.csv carry every run's result and every check's verdict.

…h checks notice a fault at the environment-trainer boundary: the status rule, a strict curve check, byte-identical weights, or reconciliation with an env-side record

21 client faults (every field that crosses, with zero/stale/swap/f16/clip, plus episode-end and step faults), 2 trainer-log faults, 2 the bridge rejects, 3 controls; doses from one cell to every cell; Pendulum (traced, with a first-difference localiser) and HalfCheetah; 590 runs.
… always; byte-identical weights or reconciliation with an env-side record catches every one

590 registered runs, SB3 PPO on Pendulum and HalfCheetah through the bridge, 21 client faults and 2 trainer-log faults at doses from one value to every value. Every check and prediction holds: the status rule caught 0-4% at doses up to 10% and 39% at every value; weights caught all 504 client-fault runs and no control; the record caught what changes the books; weights or record caught all 546 silent runs; the first difference located all 351 traced runs.
@tactino
tactino merged commit 7cdd1cf into main Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant