A dead-letter queue holds integration messages that couldn't be processed after their allowed retries, so one bad payload doesn't block the pipeline or disappear without trace. Retry temporary errors such as timeouts, 429s and 5xx responses with backoff. Send permanent errors such as validation failures straight to the DLQ, alert on them, and replay them once fixed.
What a dead-letter queue is and when to use it
In a queue-based integration, a worker takes a message, processes it, and deletes it on success. When processing keeps failing, the message has two bad options. It can stay in the main queue, where it's retried indefinitely, wastes capacity and can hold up the messages behind it in ordered queues. Or it can be dropped, and the data is lost. A dead-letter queue is a third option: a separate holding area where failed messages wait, with their error context, for someone to look at them.
Most messaging services support this directly. Amazon SQS, for example, moves a message to a configured DLQ once it has been received more times than the maxReceiveCount in the queue's redrive policy, and supports redriving messages back to the source queue. Anypoint MQ and other brokers offer similar features. In Salesforce-native designs, a custom object such as Integration_Error__c often plays the DLQ's role, holding the payload reference, error and status for each failed message.
Use a DLQ when:
- The integration is asynchronous and someone needs to know if a message never arrives.
- The data matters financially or operationally: orders, invoices, payments, entitlements, such as in a Salesforce to NetSuite order-to-cash sync.
- Some failures can only be fixed by a person, such as a missing product mapping, an inactive user or a validation rule.
- You need to show what failed, when, and how it was resolved, for audit or for a partner.
For a synchronous call where the user sees the error straight away, a DLQ adds less, although logging still helps.
Retry vs dead-letter rules
Most of what makes a DLQ useful comes from classifying errors correctly. The basic rule is to retry what might succeed on its own and dead-letter what won't.
| Error | Typical source | Action |
|---|---|---|
| Timeout, connection reset | Network, overloaded endpoint | Retry with backoff |
| HTTP 429 Too Many Requests | Rate limit | Retry after the Retry-After header if present, otherwise back off |
| HTTP 500, 502, 503, 504 | Temporary server-side failure | Retry with backoff |
UNABLE_TO_LOCK_ROW |
Salesforce record lock contention | Retry with backoff; reduce parallelism if it keeps happening |
REQUEST_LIMIT_EXCEEDED |
Salesforce API allocation used up | Pause the pipeline and alert; retrying immediately makes it worse |
| HTTP 401 | Expired token | Refresh the token once and retry; if it still fails, pause and alert |
HTTP 400 / 422, FIELD_CUSTOM_VALIDATION_EXCEPTION, REQUIRED_FIELD_MISSING, INVALID_FIELD |
Bad data or mapping | Dead-letter immediately |
HTTP 403, INSUFFICIENT_ACCESS_OR_READONLY |
Permissions | Dead-letter and alert; it's a configuration issue |
DUPLICATE_VALUE on an idempotency key |
Message already processed | Treat as success; don't dead-letter |
| Schema or parse error | Malformed payload ("poison message") | Dead-letter immediately |
Retry with backoff and jitter. Spread retries out so a struggling downstream system has time to recover, and so parallel workers don't all retry at the same moment:
function nextDelaySeconds(attempt):
base = 2
cap = 900 // 15 minutes
exp = min(cap, base * 2^attempt) // 2, 4, 8, 16 ... capped
return random(0, exp) // "full jitter"
function process(message):
try:
handle(message) // must be idempotent
ack(message)
catch TransientError as e:
if message.attempts < MAX_ATTEMPTS: // e.g. 5-8
requeue(message, delay = nextDelaySeconds(message.attempts))
else:
deadLetter(message, e, reason = "RETRIES_EXHAUSTED")
catch PermanentError as e:
deadLetter(message, e, reason = "NON_RETRYABLE")
In Apex, System.enqueueJob(job, delayInMinutes) lets a Queueable schedule its own retry with a delay of up to 10 minutes. For longer schedules, store the next retry time on the error record and use a scheduled job to pick it up.
Retries are only safe if processing is idempotent. Otherwise every retry is another chance of a duplicate. See designing idempotent ingestion layers.
Alerting and replay workflow
A DLQ that nobody looks at just loses data more slowly. Set up a workflow around it.
Alert on signals, not noise:
- New permanent failures. Alert on any message dead-lettered for a non-retryable reason, grouped by error type so 300 identical validation errors arrive as one alert.
- DLQ depth and age. Alert when the DLQ has more than a set number of messages, or when the oldest one is older than your agreed triage window.
- Retry spikes. A sudden rise in retries often comes before an outage or a limit being exhausted.
- Silence. No messages processed during a period when you'd normally see traffic usually means something upstream has broken.
- Platform headroom. Daily API usage approaching its allocation, read from the Salesforce REST
/limitsresource.
Replay workflow:
- Triage. The on-call owner groups DLQ messages by error code and finds the root cause: a data issue, a mapping gap, a code defect or a configuration change.
- Fix the cause first. Correct the data, add the missing mapping or deploy the fix. Replaying before the cause is fixed just fills the DLQ again.
- Replay in a controlled way. Move messages back to the main queue (for example with SQS redrive) or re-run them with a replay script, starting with a small sample before the rest.
- Verify. Confirm the target records are correct in Salesforce or the downstream system.
- Close out. Mark each message as resolved, replayed or discarded, with who did it and why. Discarding should be deliberate and recorded.
Automatic or manual replay? Automatic redrive suits failures with a known temporary cause, such as replaying everything after a confirmed downstream outage ends. Manual or approved replay suits permanent errors, because someone has to fix the cause first. Many teams use both: automatic for "retries exhausted during an incident", manual for "non-retryable".
Logging fields to keep
Log enough to diagnose and replay without opening the raw payload, and store as little personal data as you can. That matters for UK GDPR and EU GDPR. For each message, consider keeping:
| Field | Purpose |
|---|---|
correlation_id |
Traces one business transaction across every system and step |
message_id / idempotency_key |
Makes replays safe and duplicates detectable |
source_system, target_system |
Shows where it came from and where it was going |
operation, entity_type |
For example "upsert", "Invoice" |
external_id, salesforce_record_id |
Links to the business record in each system |
attempt_count, first_attempt_at, last_attempt_at |
Retry history and age |
error_category |
Transient or permanent, for routing and reporting |
http_status, error_code, error_message |
The exact failure; truncate long messages |
payload_ref or hash |
A pointer to the stored payload (encrypted, access-controlled) instead of the raw body in logs |
schema_version |
Shows which mapping version processed it |
status, resolved_by, resolved_at, resolution_note |
Audit trail for replay or discard |
Set retention deliberately. Keep logs and DLQ payloads long enough for triage and audit, and no longer than your data protection policy allows.