Make background workers safe across control-plane replicas - #64
gciavarrini wants to merge 8 commits into
Conversation
DB-backed claiming with optimistic updates on SQLite and SKIP LOCKED on Postgres. Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Second control-plane replica in compose. assert one cleanup delete publish per scheduler tick across replicas. Fail fast when the second replica is not reachable. Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
PR Summary by QodoLease deferred deletions across control-plane replicas
AI Description
Diagram
High-Level Assessment
Files changed (9)
|
Code Review by Qodo
1.
|
Use compose profile ha so test-up stays single-replica. HA spec brings cp2 up on demand, polls health, and stops it after. Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Do not clear deletion_claimed_until when recording a publish attempt. Release claims on publish failure, lookup errors, and cycle cancel so retries are not blocked for the full TTL Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
| "github.com/dcm-project/control-plane/internal/sp/store/model" | ||
| ) | ||
|
|
||
| // deletionClaimTTL keeps a claimed deletion off other replicas while this |
There was a problem hiding this comment.
The five-minute deletion lease remains the only duplicate-dispatch guard after a successful publish, but it is not renewed while the asynchronous agent deletion is in flight. If the agent takes longer than five minutes to acknowledge, another replica can reclaim and republish the deletion. Each Publisher.publish call generates a fresh JetStream message ID, so the later publish is not deduplicated. Please renew the lease until acknowledgment or use a stable operation identifier/idempotent dispatch mechanism.
After a successful delete publish, extend the claim for the await-ack window. Use a stable Nats-Msg-Id and matching JetStream dedup window so a second replica does not deliver a duplicate delete while the first is still in flight. Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Resolve SP subsystem compose conflict: keep shared .env from main and the HA control-plane-2 profile from this branch. Assisted-By: Claude (Anthropic) Signed-off-by: Gloria Ciavarrini <gciavarrini@redhat.com>
Make deferred-deletion cleanup safe when multiple control-plane instances share the same Postgres database.
Each replica leases
SCHEDULEDrows before publishing to the agent, so only one instance drives a given deletion at a time.deletion_claimed_untilandClaimPendingDeletions(optimistic onSQLite,
FOR UPDATE SKIP LOCKEDon Postgres)control-plane-2in composeProvider health-check claiming is out of scope here:
mainuses the agent heartbeat monitor (MarkStaleUnavailable), which is already a single atomic update.Fixes
https://redhat.atlassian.net/browse/FLPATH-4630