Skip to content

[Bug] Pipeline worker no-progress can leave jobs wedged indefinitely #4342

Description

@sachintandon

Summary

A background pipeline worker can stop advancing a job while the server remains available. The job last_processed_at then remains stale, there is no useful per-worker heartbeat to distinguish a worker stall from normal queueing, and recovery may depend on a server process restart. The underlying stall trigger is not yet established.

Expected behavior

  • Expose worker/task heartbeat and pending-work/progress metrics so a stalled pipeline is observable.
  • Ensure per-item exceptions record an error and advance retry/progress metadata before the item is retried or backed off.
  • Do not silently stop processing a job after a transient exception.

Please add regression tests for a failing item followed by successful retry and for stale-worker detection. Related historical reports include #825 and #3886, but this issue focuses on persistent worker no-progress and observability.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions