Skip to content

app_lifecycle: revive always-on apps whatever their stored status - #240

Open
ClaydeCode wants to merge 5 commits into
mainfrom
fix/clayde/always-on-revive-running
Open

ClaydeCode wants to merge 5 commits into
mainfrom
fix/clayde/always-on-revive-running

Conversation

@ClaydeCode

Copy link
Copy Markdown
Contributor

Fixes #239.

When shard_core is killed before its shutdown has run compose down on every app, an always_on app keeps a RUNNING row while its containers are gone. The control tick skipped RUNNING always-on apps, so such an app stayed down until someone opened it.

The tick now hands every always-on app to start_app (size check kept). start_app already dispatches on the real container state, skips non-revivable statuses, and is a throttled no-op for a stack that is really up. The cost is one docker compose ps plus one docker inspect per always-on app per tick. The public app store has two always-on apps (mosquitto, syncthing).

Changes

  • app_lifecycle._control_app_time: drop the status != RUNNING guard from the always-on branch.
  • Tests: the always-on tick test now covers PAUSED, STOPPED, DOWN and RUNNING. A new test covers the size gate. The start_app test covers a RUNNING row whose stack is missing or exited.
  • agents.md: documents how the control tick revives always-on apps.

Verification

  • The full suite passed locally (393 tests) on the fix commit. Later commits only change tests and docs, and the touched test files were re-run green (62 tests).
  • Mutation checks: restoring the old guard fails the [running] case, and dropping the size check fails the size-gate test. Removing RUNNING from _REVIVABLE_STATUS, or making start_app return early on RUNNING, fails the start_app test.

Review panel

  • Adversarial reviewer: ran. No blocking findings.
    • Advisory: the agents.md wording overstated the behaviour (fixed). The exact RUNNING-and-missing-stack case had no start_app-level test (fixed).
    • Advisory, not acted on: an always-on app that ships a one-shot container that exits by design would now be re-uped on every tick. No always-on app in the store does this today (mosquitto and syncthing each run a single service with restart: always). Worth keeping in mind when app-repository guidance for always-on apps is written.
  • Test adversary: ran. No blocking findings. It broke the code 7 ways and restored every file with git checkout.
    • Three mutations survived: dropping the size check, dropping PAUSED/DOWN from the revive, and skipping a RUNNING row whose containers exited. Each now has a covering test.
  • DevEx / readability: ran. No blocking findings.
    • Fixed: the test was renamed to what it pins. The size-gate test now passes only because of the gate. The agents.md sentence now names the disk and size exceptions and says it is independent of the pause flag. The two agents.md commits were squashed.
    • Kept, against its suggestion: the (RUNNING, "exited") case. The issue's scenario can leave containers stopped rather than removed, for example after a host reboot or a controller cutover.
  • Not dispatched, because the diff has no auth, input parsing, schema, SQL, route shape or Sundial UI: security, DB/migration, API contract, UX.

🤖 Generated with Claude Code

ClaydeCode and others added 5 commits September 23, 2026 12:25
When shard_core is killed before its shutdown has run compose down on every
app, an always-on app keeps a RUNNING row while its containers are gone. The
control tick skipped RUNNING apps, so such an app stayed down until someone
opened it. start_app already dispatches on the real container state and is a
throttled no-op for a stack that is really up, so the tick now always asks it.

Fixes #239

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This is the exact state issue #239 leaves behind; the tick-side test mocks
start_app, so the revive path itself needs its own case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin every revivable stored status (DOWN is the one a clean shutdown leaves),
the size check that the fix deliberately keeps, and a RUNNING row whose
containers exited rather than vanished (a host reboot or the controller
cutover stopping leftovers). Each case failed under a matching mutation of
the production code in review. The rationale comment moves out of the test
and into this history.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Make the size-check case pass only because of the size gate: a DOWN, long-idle
app would otherwise be started or fall through to idle stop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@ClaydeCode
ClaydeCode requested a review from max-tet September 23, 2026 12:58
@ClaydeCode

Copy link
Copy Markdown
Contributor Author

CI: unit tests, ruff and build are green. Backend types drift check is red, and the cause is outside this branch. freeshard-controller main has moved on since the last resync (#238): its vendored shard_model now has a CUTTING_OVER shard status and a cloud field on ShardUpdateDb, probably from the cutover work in FreeshardBase/freeshard-controller#495. Any PR opened now will show this failure until a separate just get-types resync lands. The check is not required by the protect main ruleset, so I left it out of this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An always-on app stays down after shard_core is killed mid-shutdown

1 participant