Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions docs/spec/SPEC-011-daemon-architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,6 +210,13 @@ The daemon writes its PID to `~/.netclaw/netclaw.pid` for lifecycle management.
shutdown. The daemon handles SIGTERM by draining active sessions and stopping
the actor system cleanly.

The session journals each accepted input before it acknowledges the source.
The record retains the text, media, source message ID, and original authority.
A completed reply, a started tool batch, or a terminal failure consumes the
input ID. The actor restores unconsumed records from the journal after a cold
start. A retry with the same stable source message ID does not add a second
record. A source without a stable ID cannot use this deduplication rule.

During drain, a session can stop a tool task that waits only for durable
approval prompts. The session waits for the tool task to stop before it
acknowledges drain. Its journal retains the prompts and completed sibling
Expand Down
2 changes: 2 additions & 0 deletions openspec/changes/resume-interrupted-sessions/.openspec.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-18
104 changes: 104 additions & 0 deletions openspec/changes/resume-interrupted-sessions/design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
## Context

See [proposal.md](proposal.md) for the user outcome. The actor currently acknowledges a model-only request and each buffered request before a journal event stores that input. `TurnRecorded` stores one user message after the reply. A stop clears the actor-local queue. `RestartRecoveryService` warms prior sessions but does not start a model call. The manifest exists only for a configuration restart.

The [engineering glossary](../../../docs/spec/GLOSSARY.md) defines shared session terms. `TurnContextRecord` already stores the authority fields that a recovered tool approval uses. The channel bindings own output subscribers. A warm session alone cannot deliver a recovered reply to Slack, Discord, or Mattermost.

## Goals / Non-Goals

**Goals:**

- A journal record precedes each accepted input acknowledgment.
- The journal and snapshots retain accepted input until a durable terminal event consumes it.
- A graceful stop confirms a model task stop before it grants a resume candidate.
- A short restart resumes only eligible work under its original authority.
- One model call receives the ordered accepted queue after the original turn ends.

**Non-Goals:**

- A crash does not create an automatic resume candidate.
- A tool with an uncertain effect does not run again without new user input.
- A new reminder definition does not represent interrupted work.
- A channel without a confirmed output route does not get an automatic model call.

## Decisions

### D1. The journal owns admitted input

Add `InputAdmitted` with an input ID, a stable source message ID when one exists, text, media, executable text, source IDs, received time, and a `TurnContextRecord`. The actor persists it before the ack. The record excludes actor refs and raw `MessageSource`. Reject a missing or invalid authority record before the model call. The actor uses a source ID only within its channel and session scope for deduplication. Sources without a stable ID cannot claim retry deduplication.

The actor keeps admitted records in an ordered state list. `TurnRecorded` identifies every input in its model call. `ToolBatchStarted` identifies the inputs before any tool runs. A durable terminal failure event consumes a model-only input. This prevents a failed turn from starting again after a short stop. The live callback and journal replay use the same input IDs. The snapshot stores pending records and a bounded recent source-ID ledger. The actor skips a snapshot while a pending input also appears only in transient history.

The old `TurnRecorded.UserMessage` remains for journal compatibility. Replay uses admitted records when the new consumed-ID list exists. Replay uses the old field for earlier journal records. This keeps old sessions readable.

Alternative: Store the input only in a stop manifest. That misses a stop after an ack and before manifest creation. It also duplicates session authority outside the journal.

### D2. Drain returns a classified result

`PrepareForDaemonRestart` asks the actor to stop new work. A live model call gets a short grace. If it completes, the actor handles its normal result. If it remains active, the actor cancels its token and awaits the exact `SessionLlmInvoker` task. A call ID rejects stale task results. The actor grants a candidate only when the task stopped, no tool batch started for that turn, and no user-visible text escaped. A confirmed queue can form a candidate after a completed reply. Approval-only work follows the existing durable approval path without a candidate.

The actor returns a typed drain result with the candidate input IDs or a blocked reason. `SessionDrainHelper` collects results. The daemon writes one manifest after all drain replies. Both the config restart and normal coordinated stop use this path. The manifest write is atomic. It stores an absolute ten-minute deadline per candidate. A timeout or failure produces a warning and no executable candidate.

Alternative: Infer candidates from actor phase or session IDs. Phase does not prove task cancellation, and a warm session can contain a completed turn.

### D3. Recovery checks the manifest and the session journal

The recovery service reconciles the session catalog and warms listed sessions. After channel services start, it prepares the output route for a candidate. It then sends an internal resume request with the candidate IDs and absolute deadline. The actor checks the deadline again. It checks that the IDs still match its pending journal records, no newer turn started, and the recorded authority is valid. The actor resumes the original turn first. It then sends one ordered follow-up model call for accepted queued input.

The service uses existing channel gateway `StartProactiveThread` messages to rebuild Slack, Discord, and Mattermost bindings. Those messages must confirm the output subscriber before the resume request. A TUI or SignalR session needs a live attachment; otherwise the candidate stays blocked and the service reports it. The service must retain a candidate until it succeeds or expires. It does not reset the deadline after another process start.

Alternative: Schedule a generic reminder. That creates a new turn, can abandon a parked approval, and may run under authority derived from the reminder rather than the original input.

### D4. Authority and output safety gate

Each admitted record carries its original `TurnContextRecord`. The actor reconstructs it with `TurnContext.TryFromRecord`. It never derives a recovery authority from a session ID. A queue with incompatible boundaries remains durable but does not auto-resume. A queue with compatible authority uses the narrowest audience and never widens tool access. The actor rejects a candidate after partial text output because the user might have seen that text. It also rejects a candidate after a tool batch starts because a tool might have an external effect.

Positive example: A Slack user sends one request. The model call stops before text or tools. The next start restores the Slack binding and resumes that input under its stored personal boundary.

Negative example: A shell tool starts before a stop. The daemon reports the pending work and does not replay the shell call.

### D5. Delivery order and failure behavior

The candidate deadline starts at interruption. The recovery service checks it before each route attempt. The actor checks it before the model call. An expired candidate causes one warning and leaves the journal record available for forensic inspection or a new user turn. A newer input supersedes an old candidate. The manifest remains on disk until each candidate reaches a terminal recovery decision. The service records a blocked route or invalid record with a clear diagnostic.

## Ordered flow

This diagram is schematic. It omits the channel ACL and persistence callbacks.

```text
source -> session: SendUserMessage
session -> journal: InputAdmitted(input ID, authority, content)
journal -> session: persisted
session -> source: CommandAck
session -> model: original turn
stop -> session: PrepareForDaemonRestart
session -> model: cancel if grace ends
model -> session: task stopped
session -> stop: candidate(input IDs) or blocked reason
stop -> manifest: atomic write(deadline)
start -> channel: restore output binding
channel -> start: binding ready
start -> session: ResumeInterruptedTurn(input IDs, deadline)
session -> journal: validate pending input and authority
session -> model: original turn, then one queued call
```

## Risks / Trade-offs

- [Input accepted during a write failure] → The actor sends a nack and starts no model call.
- [Snapshot skips an unconsumed input] → Replay starts before that snapshot, then restores the pending record.
- [Provider ignores cancellation] → Drain reaches its existing deadline and records no candidate.
- [A reply reaches the user before cancellation] → The actor reports a blocked candidate and avoids duplicate text.
- [A tool starts before cancellation] → The actor reports a blocked candidate and avoids duplicate effects.
- [A route is absent after restart] → The service reports the candidate and leaves the agent quiet.
- [Old manifest or a second process start] → The absolute deadline and input IDs prevent a new or duplicate model call.
- [A source has no stable message ID] → The actor records a unique admission ID but cannot deduplicate a source retry.

## Migration Plan

1. Add the journal and snapshot fields with new protobuf tags and a new manifest for `InputAdmitted`.
2. Keep old event fields and old journal readers intact.
3. Add actor admission and recovery tests before enabling automatic wakeup.
4. Add the normal stop manifest and recovery request after the actor contract passes.
5. Verify a graceful stop and a short restart in an isolated daemon with a persistent home.
6. Roll back by disabling automatic wakeup in the new binary. Keep the new journal decoder so accepted input remains readable.
37 changes: 37 additions & 0 deletions openspec/changes/resume-interrupted-sessions/proposal.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
## Why

A graceful stop can acknowledge input before the journal stores it. A model call can also delay drain until the global stop limit. A short restart then leaves accepted work quiet or loses queued context.

Source: [PRD-001 FR-003 and FR-016](../../../docs/prd/PRD-001-netclaw-mvp.md).

## What Changes

- Persist each accepted input, its order, media, source ID, delivery context, and authority before the input ack.
- Record the input IDs that a completed turn consumes. Restore accepted input from the journal after cold recovery.
- Stop an eligible model call after a short grace and wait for its task to stop.
- Save one bounded resume candidate after any graceful stop. The deadline expires ten minutes after interruption.
- Resume the original turn under its recorded authority. Deliver the accepted queue in one ordered follow-up model call.
- Leave a completed reply with no accepted queue quiet. Block auto-resume for approvals, uncertain tool effects, and partial replies.
- Report blocked work to the operator.
- Correct the container contract for stop signals and the persistent state volume.

The MVP change excludes automatic crash recovery, replay of uncertain tool effects, and a new reminder schedule. It does not shorten the global stop limit.

## Capabilities

### New Capabilities

None.

### Modified Capabilities

- `session-resume`: Durable input admission and bounded restart recovery for confirmed interruptions.
- `daemon-container`: The entrypoint forwards stop signals, and the state volume retains restart data.

## Impact

This change affects the session actor, journal events, protobuf mappings, restart manifest, recovery service, and container contract. It adds no user configuration property.

### Security and operational impact

The actor restores the original authority from a durable record. It rejects an incomplete record and does not replay a tool with an uncertain effect. A stale candidate expires without an agent turn. A process stop must complete its journal and manifest writes before exit.
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
## MODIFIED Requirements

### Requirement: Image entrypoint auto-starts netclawd

The image SHALL start `tini` as PID 1. Its supervisor SHALL start `netclawd` and forward a container stop signal to it. The supervisor SHALL wait for the daemon to finish graceful drain before it exits.

#### Scenario: docker run starts the daemon

- **GIVEN** the image is present locally with valid configuration and identity files
- **WHEN** an operator starts the container
- **THEN** `tini` is PID 1 and the supervisor starts `netclawd`
- **AND** the daemon binds its HTTP port within 60 seconds

#### Scenario: Pod stop preserves a resume candidate

- **GIVEN** the container has a persistent operator state volume and an eligible interrupted session
- **WHEN** the pod sends a graceful stop signal with enough termination time
- **THEN** the supervisor forwards the signal and waits for the daemon to exit
- **AND** the state volume retains the candidate for the next container start

### Requirement: Operator state mounts at /home/netclaw/.netclaw

The image SHALL declare `VOLUME /home/netclaw/.netclaw`. The volume SHALL hold identity, configuration, session data, and restart candidates. The image SHALL not include operator credentials or identity files.

#### Scenario: Operator bind-mounts an initialized home

- **GIVEN** an operator has an initialized Netclaw home on the host
- **WHEN** they mount it at `/home/netclaw/.netclaw` and start the container
- **THEN** the daemon reads identity and configuration from that directory
- **AND** it writes session state and restart candidates to the same directory

## RENAMED Requirements

- FROM: `Operator state mounts at /root/.netclaw`
- TO: `Operator state mounts at /home/netclaw/.netclaw`
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
## ADDED Requirements

### Requirement: Accepted input survives a graceful stop

The session SHALL record each accepted input before it acknowledges that input. The record SHALL retain a stable input ID, source ID, order, text, media, authority, and delivery context. A failed journal write SHALL reject the input.

Use the [engineering glossary](../../../../../docs/spec/GLOSSARY.md) for shared terms.

#### Scenario: Input ack follows its journal record

- **GIVEN** a user sends input to a ready session or a busy session
- **WHEN** the journal confirms the accepted input record
- **THEN** the session acknowledges the input
- **AND** cold recovery restores its text, media, order, and original authority

#### Scenario: Failed journal write rejects input

- **GIVEN** the journal cannot store an input record
- **WHEN** the user sends that input
- **THEN** the session rejects the input with a visible error
- **AND** the session does not start a model call for it

#### Scenario: Source retry does not duplicate accepted input

- **GIVEN** the journal stores an input with a stable source message ID
- **WHEN** the source retries that same message after it loses the ack
- **THEN** the session acknowledges the existing input
- **AND** the session does not add a second copy to the queue

### Requirement: Graceful stop records only safe resume candidates

The daemon SHALL record a resume candidate after any graceful stop only for confirmed interrupted model work or accepted queued input. The candidate SHALL have one absolute deadline ten minutes after interruption. A stopped model task SHALL be confirmed before the session acknowledges drain.

#### Scenario: Interrupted model call creates a candidate

- **GIVEN** an admitted turn has a live model call and no tool batch has started in that turn
- **WHEN** a graceful stop interrupts the call after a short completion grace
- **THEN** the session waits for the model task to stop
- **AND** the daemon records a candidate for that original turn

#### Scenario: Completed reply with queued input creates a candidate

- **GIVEN** a turn finishes during drain and accepted queued input remains
- **WHEN** the session acknowledges drain
- **THEN** the daemon records a candidate for the accepted queue

#### Scenario: Completed reply without queued input stays quiet

- **GIVEN** a turn finishes during drain and no accepted queued input remains
- **WHEN** the daemon starts again
- **THEN** the daemon starts no model call for that session

#### Scenario: Tool effect or partial reply blocks automatic resume

- **GIVEN** the interrupted turn has a tool batch or user-visible partial text
- **WHEN** the daemon stops
- **THEN** the daemon does not record an executable model resume candidate
- **AND** it reports the blocked work to the operator

### Requirement: Eligible restart resumes original work once

The daemon SHALL resume an eligible candidate only before its deadline and after its output route is ready. The session SHALL use the recorded authority and input IDs. It SHALL not create a new reminder turn or ask the user for stored context.

#### Scenario: Cold recovery resumes the original turn

- **GIVEN** a graceful stop confirmed model cancellation and stored the admitted input
- **WHEN** the daemon starts before the candidate deadline
- **THEN** the session resumes one model call for the original turn
- **AND** it uses the recorded requester and trust boundary

#### Scenario: Accepted queue follows the original turn

- **GIVEN** a canceled turn and multiple accepted queued messages
- **WHEN** the resumed original turn finishes
- **THEN** the session sends the queued messages in their original order
- **AND** it uses one follow-up model call for that queue

#### Scenario: Expired candidate stays quiet

- **GIVEN** the absolute candidate deadline has passed
- **WHEN** the daemon starts or tries to resume the session
- **THEN** it starts no model call from that candidate
- **AND** it reports and removes the expired candidate once

#### Scenario: Output route is absent

- **GIVEN** a candidate has no live route that can deliver the resumed output
- **WHEN** the daemon starts before its deadline
- **THEN** it does not run a model call yet
- **AND** it reports the blocked candidate to the operator

#### Scenario: Newer work supersedes a candidate

- **GIVEN** a newer user turn starts before the recovery request reaches the session
- **WHEN** the older recovery request arrives
- **THEN** the session rejects the stale request and starts no duplicate call

#### Scenario: Incompatible queued authority blocks automatic resume

- **GIVEN** accepted queued messages have incompatible trust boundaries
- **WHEN** the daemon tries to resume that queue
- **THEN** the session keeps the messages in the journal
- **AND** it reports the blocked queue instead of widening tool authority
Loading
Loading