top of page

The Retry + Dead-Letter Pattern That Saved Us From Silent Message Loss

Writer: Siddhesh Dambe
Siddhesh Dambe
16 hours ago
3 min read

We run a set of NestJS microservices that talk to each other over Azure Service Bus queues. For a while, our error handling assumption was simple and wrong: "if the consumer throws, the message just gets redelivered." What really happened, a transient downstream failure would silently drop a message.



Queue Message Drops
Queue Message Drops

Table of Contents


The failure mode


Azure Service Bus will redeliver a message if your handler throws and you don’t complete it but only up to maxDeliveryCount. After that, the message either dead-letters automatically (if DLQ is enabled) or gets discarded, depending on configuration. Without explicit backoff between attempts, a transient failure (say, a downstream service having a bad 10 seconds) can burn through all delivery attempts almost instantly, and the message dead-letters or vanishes before the dependency recovers. If nothing is watching the dead-letter queue, that failure is invisible.


The fix: backoff + DLQ + alerting


Three pieces, each solving a different part of the problem:


Exponential backoff on the consumer


Instead of retrying immediately, we track delivery count and delay reprocessing by a simple modification to the catch block

catch (err) {
	if (deliveryCount >= this.maxRetries) {    
		await message.deadLetter({
			reason: 'MaxRetriesExceeded',     
			errorDescription: err.message,
		});
		return;
	}
	const delay = this.baseDelayMs * Math.pow(2, deliveryCount - 1);
	await message.abandon(); // returns to queue for redelivery
}

The key detail: abandon() alone just makes the message immediately available again it doesn’t back off. If your failure is transient-but-not-instant (a downstream service restarting, a rate limit), you need real delay between attempts, either via scheduled enqueue time or a lock-duration-aware wait, not a tight retry loop.


Dead-letter queue for poison messages


Anything that exhausts retries goes to the DLQ instead of disappearing. This turns “silently lost” into “parked somewhere inspectable.” We tag the dead-letter reason and error description on the way in, so when we do look, we’re not reverse-engineering what happened from scratch.


Alerting on DLQ depth


This is the piece that actually closes the loop. A dead-letter queue that nobody watches is just a slower way to lose messages. We alert on DLQ message count per queue — even a single message triggers a look, since in our system dead-lettering is always a symptom of something, not expected steady-state noise.


Why this matters more now than it used to


This pattern isn't new it's standard queue hygiene, and most teams running Service Bus or SQS eventually reinvent some version of it. But it's become more load-bearing since we started putting LLM calls in the middle of these pipelines.


LLM API calls fail transiently far more often than a typical database call: rate limits, timeouts, provider-side hiccups, occasional 5xxs under load, and latency variance that turns a normal request into a de facto timeout. A pipeline that was "reliable enough" with a DB call in the hot path can start dropping messages regularly once an LLM call sits where the DB call used to be not because the LLM integration is buggy, but because transient failure is simply a much bigger share of its failure modes.


That shift is worth sizing up before it bites you. If a dependency fails transiently 0.1% of the time and you retry three times with no backoff, you'll still lose the occasional message during a real outage, but rarely enough that it stays invisible. If that dependency is an LLM call failing 2–5% of the time under normal load, the same retry policy turns from a safety net into a slow leak and without DLQ alerting, that leak looks like nothing until someone notices missing data downstream, usually well after the fact.


This fix isn't exotic: back off properly, dead-letter deliberately instead of by accident, and alert on DLQ depth as if it were a first-class error metric because functionally, it is. Do that before you bolt any action into a queue consumer, not after the first quiet data-loss incident forces the conversation.


Glossary


de facto timeout: A "de facto" timeout refers to a situation where a break, stoppage, or expiration occurs in practice or reality, even though it wasn't officially called, intended, or explicitly written into the rules.


Exponential Backoff: A retry strategy where the waiting time increases after each failed attempt, reducing pressure on an unhealthy service.


Dead-Letter Queue (DLQ): A separate queue where messages are moved after they repeatedly fail processing or cannot be successfully delivered.


Poison Message: A message that consistently causes processing failures, often because of invalid data, an unexpected format, or an application bug.


bottom of page