TIL / A queue consumer needs to catch SIGTERM, not just crash on it

A queue consumer needs to catch SIGTERM, not just crash on it

KubernetesReliabilityQueues

The problem

A rolling deploy or a KEDA scale-down sends SIGTERM to a pod, waits out terminationGracePeriodSeconds, then sends SIGKILL. A queue consumer that has no signal handler just gets killed mid-message: whatever it was processing is lost or, worse, left in a half-applied state if the work wasn’t idempotent.

The fix

Install a SIGTERM handler that flips a flag, let the current message finish, stop pulling new ones, and exit cleanly before the grace period runs out.

import signal

shutting_down = False

def handle_sigterm(signum, frame):
    global shutting_down
    shutting_down = True

signal.signal(signal.SIGTERM, handle_sigterm)

while not shutting_down:
    message = queue.receive(wait_seconds=5)
    if message:
        process(message)
        queue.delete(message)

Gotcha

The grace period has to be longer than the slowest single message can realistically take, or Kubernetes sends SIGKILL anyway before the handler finishes - check terminationGracePeriodSeconds against your actual p99 processing time, not the average.