TIL / A stale-job reaper stops a crashed worker from wedging a pipeline forever

A stale-job reaper stops a crashed worker from wedging a pipeline forever

PipelinesReliabilityPostgreSQL

The problem

A multi-stage document pipeline marks each job processing when a worker picks it up and done when it finishes. That’s fine as long as every worker finishes what it starts - but a worker that gets OOM-killed, loses its pod, or hits an unhandled exception leaves the job stuck at processing forever. Nothing else in the pipeline will ever touch it again, and it silently disappears from throughput without ever showing up as a failure.

The fix

A periodic reaper query finds jobs that have been processing past a reasonable timeout and resets them to pending (or failed, past a retry limit) so a healthy worker can pick them back up.

UPDATE jobs
SET status = 'pending', attempts = attempts + 1
WHERE status = 'processing'
  AND updated_at < now() - interval '15 minutes'
  AND attempts < 3;

Gotcha

The timeout has to be longer than the slowest legitimate job, or the reaper starts fighting a worker that’s still genuinely working - pick it from real p99 duration, not a guess, and log every reap so a job that keeps getting reset (rather than completing) shows up as a real alert instead of quietly retrying forever.