RabbitMQ
Dead-letter queues and poison message handling
A dead-letter queue (DLQ) is where messages go when consumers can’t process them. Used correctly, it’s your early-warning system and debugging buddy — not a dumping ground.
1. Common reasons messages hit the DLQ
- Deserialization failures (bad payload).
- Business rule violations (e.g., invalid state transitions).
- Permanent external errors (e.g., missing reference data).
2. Set clear DLQ policies
I usually configure queues with:
- A limited number of retries via
x-deathheader / redelivery count. - A DLX (dead-letter exchange) and DLQ per logical workload.
- TTL for DLQ messages if they’re not needed forever.
3. Attach metadata for debugging
When you send messages to the DLQ, including:
- Exception type and message (sanitized).
- Where it failed (service name, handler name).
- Correlation ID / trace ID.
4. Don’t auto-replay everything
It’s tempting to “drain the DLQ back into the main queue” after a fix. Instead, I prefer:
- Replay in small batches.
- Run replays during low-traffic windows.
- Track replay success vs still-failing messages.