Repository navigation
Fix short writes on a full PERSIST volume losing acknowledged items - #163
Merged
Merged
Conversation
…160) Root cause: FileStore ignored the byte count returned by Deno.FsFile.writeSync(). On a full volume the OS performs a short write and returns fewer bytes without an error, so a partial log line was reported as success. The next append was glued onto it and replay dropped both; the shutdown snapshot was fsynced and renamed over a complete persist.dat while truncated. - Write every byte or throw (EFBIG/ENOSPC surfaces on the retry). - Roll a failed append back to the previous end of the log (truncate and seek), so the log stays line-aligned. - A snapshot that cannot be fully written throws before the rename, keeping the old persist.dat; shutdown logs the error and exits 1. - Manager logs enqueue/dequeue before mutating memory, so a request that cannot be persisted fails (500) without changing queue state. Regression tests use `ulimit -f` to force real short writes without needing root to mount a tiny volume. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jonbaldie
added a commit
that referenced
this pull request
Oct 10, 2026
Enqueue and dequeue already write the log before changing memory (#163). A failed write therefore returns 500 and leaves the queue unchanged. Nothing covered a write that fails outright once persist.dat is already at the size cap, so reversing that order would again drop an item on a 500 dequeue and keep a rejected enqueue. The regression test fills the log to the ulimit cap, then checks that a 500 dequeue and a 500 enqueue leave length unchanged and that only the accepted item is recovered after restart.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #160
Root cause
FileStoreignored the byte count returned byDeno.FsFile.writeSync(). When the volume is full, the OS does a short write: it returns fewer bytes and no error, and only the next write fails (probe:first write 102400 of 204800, thenFile too large (os error 27)). So:200. The next event was then glued onto it, and replay dropped both.replace): the truncated temp file was fsynced and renamed over a completepersist.dat.Fix
FileStore.writeAllloops until every byte is written, so the OS error comes through.replacethrows before the rename, so the oldpersist.datstays. On shutdown,main.tslogsFailed to flush data to persist.dat: …and exits 1 instead of printingGoodbye!.Manager.enqueue/dequeuewrite to the log before changing memory. A request that can't be persisted returns500and leaves the queue unchanged.Verification
tests/short_write_test.tsruns the real server underulimit -f(with SIGXFSZ ignored). That forces real short writes and works on CI without root. Both tests failed onmain, matching the issue's two symptoms, and pass on this branch (3/3 runs).deno compileon a 2 MiB HFS+ RAM disk (real ENOSPC):journey2_diskfull.sh: the 200 KB enqueue gets500, and after SIGKILL and restart bothbefore-fullandafter-space-freedare recovered.journey2_snapshot_nearfull.sh: exit status 1,persist.datstays at 200123 bytes, and both items are recovered.deno testlocally: 374 passed. The only failure is the Stryker runner test, because Stryker isn't installed locally; CI installs it.🤖 Generated with Claude Code