app_lifecycle: revive always-on apps whatever their stored status - #240
Open
ClaydeCode wants to merge 5 commits into
Open
ClaydeCode wants to merge 5 commits into
ClaydeCode wants to merge 5 commits into
Conversation
When shard_core is killed before its shutdown has run compose down on every app, an always-on app keeps a RUNNING row while its containers are gone. The control tick skipped RUNNING apps, so such an app stayed down until someone opened it. start_app already dispatches on the real container state and is a throttled no-op for a stack that is really up, so the tick now always asks it. Fixes #239 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This is the exact state issue #239 leaves behind; the tick-side test mocks start_app, so the revive path itself needs its own case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin every revivable stored status (DOWN is the one a clean shutdown leaves), the size check that the fix deliberately keeps, and a RUNNING row whose containers exited rather than vanished (a host reboot or the controller cutover stopping leftovers). Each case failed under a matching mutation of the production code in review. The rationale comment moves out of the test and into this history. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Make the size-check case pass only because of the size gate: a DOWN, long-idle app would otherwise be started or fall through to idle stop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Author
|
CI: unit tests, ruff and build are green. |
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #239.
When shard_core is killed before its shutdown has run
compose downon every app, analways_onapp keeps aRUNNINGrow while its containers are gone. The control tick skippedRUNNINGalways-on apps, so such an app stayed down until someone opened it.The tick now hands every always-on app to
start_app(size check kept).start_appalready dispatches on the real container state, skips non-revivable statuses, and is a throttled no-op for a stack that is really up. The cost is onedocker compose psplus onedocker inspectper always-on app per tick. The public app store has two always-on apps (mosquitto, syncthing).Changes
app_lifecycle._control_app_time: drop thestatus != RUNNINGguard from the always-on branch.start_apptest covers a RUNNING row whose stack is missing or exited.agents.md: documents how the control tick revives always-on apps.Verification
[running]case, and dropping the size check fails the size-gate test. Removing RUNNING from_REVIVABLE_STATUS, or makingstart_appreturn early on RUNNING, fails thestart_apptest.Review panel
agents.mdwording overstated the behaviour (fixed). The exact RUNNING-and-missing-stack case had nostart_app-level test (fixed).uped on every tick. No always-on app in the store does this today (mosquitto and syncthing each run a single service withrestart: always). Worth keeping in mind when app-repository guidance for always-on apps is written.git checkout.agents.mdsentence now names the disk and size exceptions and says it is independent of the pause flag. The twoagents.mdcommits were squashed.(RUNNING, "exited")case. The issue's scenario can leave containers stopped rather than removed, for example after a host reboot or a controller cutover.🤖 Generated with Claude Code