From fb456ff1b52abdc8941d231036ad49eac110e901 Mon Sep 17 00:00:00 2001 From: Hannes Fostie Date: Wed, 7 Oct 2026 13:00:30 +0200 Subject: [PATCH] Release in-flight jobs when the async supervisor's shutdown times out When a job outlives shutdown_timeout during graceful termination, the async supervisor ends with exit!, which skips the shutdown callbacks: neither the supervisor nor its threads deregister, and their claimed jobs stay claimed until a later supervisor prunes the stale registrations and fails them with ProcessPrunedError. On a TERM, which is what deploys send, the job is failed instead of going back to its queue as it does in fork mode. Worse, the supervisor and each worker wait for that same timeout, the supervisor starting slightly earlier, so it usually runs out first and calls exit! while the worker is still inside its own deregistration transaction, rolling back the release it had started. When the timeout runs out with threads still alive, deregister the supervisor before exit!, as the fork supervisor does after quitting its forks: the threads still running are deregistered with it and their claimed jobs go back to their queues. Deregister from a fresh copy of the registration, since the threads share theirs with the supervisor's in memory and a thread that already deregistered leaves a frozen record behind. QUIT is unchanged: it still exits without cleanup, as the async lifecycle test documents. The async lifecycle test for TERM past the timeout now expects the in-flight job back in its queue and a clean termination, like its fork counterpart, instead of accepting either outcome. Co-Authored-By: Claude Opus 5.5 (1M context) --- lib/solid_queue/async_supervisor.rb | 14 ++++++++++++++ test/integration/async_processes_lifecycle_test.rb | 13 +++++-------- 2 files changed, 19 insertions(+), 8 deletions(-) diff --git a/lib/solid_queue/async_supervisor.rb b/lib/solid_queue/async_supervisor.rb index f6ab38342..c22af0144 100644 --- a/lib/solid_queue/async_supervisor.rb +++ b/lib/solid_queue/async_supervisor.rb @@ -38,12 +38,26 @@ def perform_graceful_termination process_instances.values.each(&:stop) Timer.wait_until(SolidQueue.shutdown_timeout, -> { all_processes_terminated? }) + + # The exit! that follows skips the shutdown callbacks, so release the jobs + # of the threads still running now instead of leaving them to be pruned + attempt_to_deregister unless all_processes_terminated? end def perform_immediate_termination exit! end + def attempt_to_deregister + wrap_in_app_executor do + # Not our own copy: it shares the threads' registrations in memory, and + # those of threads that already deregistered are destroyed and frozen + SolidQueue::Process.find_by(id: process_id)&.deregister + end + rescue StandardError => error + handle_thread_error(error) + end + def all_processes_terminated? process_instances.values.none?(&:alive?) end diff --git a/test/integration/async_processes_lifecycle_test.rb b/test/integration/async_processes_lifecycle_test.rb index dc8e4c7ee..8185835d8 100644 --- a/test/integration/async_processes_lifecycle_test.rb +++ b/test/integration/async_processes_lifecycle_test.rb @@ -173,14 +173,11 @@ class AsyncProcessesLifecycleTest < ActiveSupport::TestCase "registered processes: #{SolidQueue::Process.all.map { |p| { id: p.id, kind: p.kind, pid: p.pid, last_heartbeat_at: p.last_heartbeat_at } }.inspect}" end - # After shutdown, the pause job may be either: - # - claimed (exit! called, no cleanup) OR - # - ready (graceful exit, job released back to queue) - # Both are valid outcomes depending on the timing race between supervisor and worker timeouts. - skip_active_record_query_cache do - job = SolidQueue::Job.find_by(active_job_id: pause.job_id) - assert job.claimed? || job.ready?, "Expected pause job to be claimed or ready, but was neither" - end + # Workers were shutdown without a chance to terminate orderly, but + # since they're linked to the supervisor, the supervisor deregistering + # also deregistered them and released claimed jobs + assert_job_status(pause, :ready) + assert_clean_termination end test "process some jobs that raise errors" do