Skip to content

Windows: peers.py inbox read crashes on os.O_NONBLOCK (bare use; defensive getattr form already used elsewhere) — and the read receipt is persisted before the crash #5511

Description

@AronSwan

Summary

On Windows, peers.py's inbox read path crashes unconditionally whenever a peer request brief carries inputs, because input_readiness() uses os.O_NONBLOCK — an attribute that does not exist on Windows. Two properties make this worse than a plain crash:

  1. It is on the main read path (read_inbox() → input_readiness(), peers.py), triggered by any brief with staged inputs — including the brief shapes in this repo's own test fixtures.
  2. The read receipt is persisted before the crash. After the AttributeError, an after-the-fact look at reads/<request_id>.json shows a recorded read (kind=receiver_cli_read, read_at set) even though the read actually failed — the receipt systematically misreports the outcome.

A one-line fix matches the pattern already used elsewhere in this same codebase (see below).

Repro (Windows, current main)

git rev-parse HEAD          # d74554b48...
python -c "import os; print(hasattr(os, 'O_NONBLOCK'))"   # False on win32

Then call the read path with a brief that has inputs (sandbox registry + workspace, minimal peers.request(...) followed by read_inbox(...)):

File "...\loopx\control_plane\collaboration\peers.py", line 547, in input_readiness
    os.open(path, os.O_RDONLY | os.O_NONBLOCK)
AttributeError: module 'os' has no attribute 'O_NONBLOCK'
  • Crash frame chain: read_inbox → input_readiness (line 547 on current main; 523 on v1.2.2).
  • Before raising, the receipt reads/<request_id>.json is already on disk with kind=receiver_cli_read and read_at written.
  • Briefs without an inputs key do not hit the path (the staging guard returns early) — the failure is deterministic for any input-bearing brief on Windows, and retrying cannot succeed.

The fix already exists in this repo — twice

The defensive form is used elsewhere; only peers.py is bare (both verified on current main, git grep -n "O_NONBLOCK" upstream/main):

  • loopx/collaboration_mcp.py:867 — getattr(os, "O_NONBLOCK", 0)
  • loopx/control_plane/todos/completion_result.py:32 — getattr(os, "O_NONBLOCK", 0)
  • vs. loopx/control_plane/collaboration/peers.py:547 — bare os.O_NONBLOCK

Suggested fix for peers.py (one line, consistent with those two precedents):

os.open(path, os.O_RDONLY | getattr(os, "O_NONBLOCK", 0))

On POSIX this is behavior-identical; on Windows the readiness probe degrades to a blocking open (or can be skipped, matching whatever the other call sites intend).

Impact notes

  • Platform: Windows (any brief with inputs); POSIX unaffected.
  • The persisted-but-failed receipt is the nastier half: any later audit of reads/ concludes the read succeeded. If there is a preferred outcome (e.g. record the read only after readiness succeeds, or record a failure kind), that would fix the misreport too — we're flagging it rather than prescribing, since the receipt-ordering may be intentional for crash recovery.

Our context

We hit this while operating a research squad on Windows against dsh 0.2.0rc2: three parallel workers all died on read_inbox with empty stdout/stderr, and the leftover receipts initially made it look like the reads had succeeded. Happy to turn the fix into a PR if the approach above is agreed — or an alternative if maintainers prefer dropping the readiness probe on Windows.

Activity

  1. added a commit that references this issue on Oct 3, 2026
  2. added a commit that references this issue on Oct 3, 2026
    6fc15f7
  3. loopx-agent commented on Oct 5, 2026

    @loopx-agent
    Collaborator

    The guarded Windows peer-file fix is already shipped in LoopX 1.2.4 via #5456. Published DSH plugin beta.6 now enforces the same 1.2.4 minimum across initialization, Driver and GoalBar, upgrading ordinary outdated installs or rejecting an owner-pinned outdated CLI before business calls.

    The plugin's Linux/Windows package-name install/remove CI passed, and released CLI 1.2.4 bootstrap and real macOS control were verified. These are not a Windows inbox/receipt or desktop-market readback. Keep that exact Windows peer brief with inputs, its receipt behavior and the full desktop journey as a platform qualification gap under #5208; no full Windows acceptance is claimed.

  4. JasonBuildAI commented on Oct 10, 2026

    @JasonBuildAI
    Contributor

    Coverage follow-up to the report, not a re-open of the crash fix: the guarded flag form you proposed and #5456 landed are in place, and the three Win32-marked regressions in tests/test_peer_collaboration.py (long store paths, binary CRLF/Ctrl-Z digests, long workspace inputs behind a junction) pin exactly the hazards here. They were skipped on the ubuntu shards by platform marker and the windows-powershell job never ran that file, so they had executed on no runner.

    PR #6073 runs those three rows in the Windows job. The receipt half of your report was accurate for 1.2.4: read_inbox committed the read before response enrichment, and 025be07 moved that commit after response validation on 2026-10-04. Nothing has held the new ordering since, so #6073 adds a test that a raising readiness probe leaves reads/ untouched and that the next completed read records it.

    Verified on a native Windows host: the three rows pass, and the new receipt test fails when record_read is moved ahead of readiness. Your remaining Windows peer/desktop platform qualification observation is untouched by this and stays with its existing owner.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions