async-job-processor-redis currently updates processing heartbeats from an Async task in the same cooperative reactor that is also running jobs.
That works for fully cooperative workloads, but it breaks for blocking job bodies. If a job blocks the reactor for long enough, the processing heartbeat expires and another worker can requeue the still-running job as abandoned.
That happened for me through the ActiveJob adapter with a blocking perform_now, and it caused duplicate execution of the same in-flight job.
I think async-job should continue to work correctly regardless of workload. Under blocking workloads it is acceptable for throughput to degrade, but it should not falsely conclude that a still-running job is abandoned and start executing it twice.
I think the processor should support:
- Running the processing heartbeat / abandoned-job recovery loop in a dedicated preemptive thread with its own Redis connection.
- Configurable timing for:
- heartbeat interval
- expiry / expiry factor
I am not arguing against heartbeat-based recovery. The issue is that job ownership is currently tied to reactor responsiveness rather than actual process liveness.
If useful, I can prepare a PR.
async-job-processor-rediscurrently updates processing heartbeats from anAsynctask in the same cooperative reactor that is also running jobs.That works for fully cooperative workloads, but it breaks for blocking job bodies. If a job blocks the reactor for long enough, the processing heartbeat expires and another worker can requeue the still-running job as abandoned.
That happened for me through the ActiveJob adapter with a blocking
perform_now, and it caused duplicate execution of the same in-flight job.I think
async-jobshould continue to work correctly regardless of workload. Under blocking workloads it is acceptable for throughput to degrade, but it should not falsely conclude that a still-running job is abandoned and start executing it twice.I think the processor should support:
I am not arguing against heartbeat-based recovery. The issue is that job ownership is currently tied to reactor responsiveness rather than actual process liveness.
If useful, I can prepare a PR.