Summary
A background pipeline worker can stop advancing a job while the server remains available. The job last_processed_at then remains stale, there is no useful per-worker heartbeat to distinguish a worker stall from normal queueing, and recovery may depend on a server process restart. The underlying stall trigger is not yet established.
Expected behavior
- Expose worker/task heartbeat and pending-work/progress metrics so a stalled pipeline is observable.
- Ensure per-item exceptions record an error and advance retry/progress metadata before the item is retried or backed off.
- Do not silently stop processing a job after a transient exception.
Please add regression tests for a failing item followed by successful retry and for stale-worker detection. Related historical reports include #825 and #3886, but this issue focuses on persistent worker no-progress and observability.
Summary
A background pipeline worker can stop advancing a job while the server remains available. The job
last_processed_atthen remains stale, there is no useful per-worker heartbeat to distinguish a worker stall from normal queueing, and recovery may depend on a server process restart. The underlying stall trigger is not yet established.Expected behavior
Please add regression tests for a failing item followed by successful retry and for stale-worker detection. Related historical reports include #825 and #3886, but this issue focuses on persistent worker no-progress and observability.