Resilience Engineering

Dead-Letter Queues (DLQ) & Error Recovery in Salesforce Integration Pipelines

By Waleed Rafique, Salesforce Certified Developer · Published

A dead-letter queue holds integration messages that couldn't be processed after their allowed retries, so one bad payload doesn't block the pipeline or disappear without trace. Retry temporary errors such as timeouts, 429s and 5xx responses with backoff. Send permanent errors such as validation failures straight to the DLQ, alert on them, and replay them once fixed.

What a dead-letter queue is and when to use it

In a queue-based integration, a worker takes a message, processes it, and deletes it on success. When processing keeps failing, the message has two bad options. It can stay in the main queue, where it's retried indefinitely, wastes capacity and can hold up the messages behind it in ordered queues. Or it can be dropped, and the data is lost. A dead-letter queue is a third option: a separate holding area where failed messages wait, with their error context, for someone to look at them.

Most messaging services support this directly. Amazon SQS, for example, moves a message to a configured DLQ once it has been received more times than the maxReceiveCount in the queue's redrive policy, and supports redriving messages back to the source queue. Anypoint MQ and other brokers offer similar features. In Salesforce-native designs, a custom object such as Integration_Error__c often plays the DLQ's role, holding the payload reference, error and status for each failed message.

Use a DLQ when:

  • The integration is asynchronous and someone needs to know if a message never arrives.
  • The data matters financially or operationally: orders, invoices, payments, entitlements, such as in a Salesforce to NetSuite order-to-cash sync.
  • Some failures can only be fixed by a person, such as a missing product mapping, an inactive user or a validation rule.
  • You need to show what failed, when, and how it was resolved, for audit or for a partner.

For a synchronous call where the user sees the error straight away, a DLQ adds less, although logging still helps.

Retry vs dead-letter rules

Most of what makes a DLQ useful comes from classifying errors correctly. The basic rule is to retry what might succeed on its own and dead-letter what won't.

Error Typical source Action
Timeout, connection reset Network, overloaded endpoint Retry with backoff
HTTP 429 Too Many Requests Rate limit Retry after the Retry-After header if present, otherwise back off
HTTP 500, 502, 503, 504 Temporary server-side failure Retry with backoff
UNABLE_TO_LOCK_ROW Salesforce record lock contention Retry with backoff; reduce parallelism if it keeps happening
REQUEST_LIMIT_EXCEEDED Salesforce API allocation used up Pause the pipeline and alert; retrying immediately makes it worse
HTTP 401 Expired token Refresh the token once and retry; if it still fails, pause and alert
HTTP 400 / 422, FIELD_CUSTOM_VALIDATION_EXCEPTION, REQUIRED_FIELD_MISSING, INVALID_FIELD Bad data or mapping Dead-letter immediately
HTTP 403, INSUFFICIENT_ACCESS_OR_READONLY Permissions Dead-letter and alert; it's a configuration issue
DUPLICATE_VALUE on an idempotency key Message already processed Treat as success; don't dead-letter
Schema or parse error Malformed payload ("poison message") Dead-letter immediately

Retry with backoff and jitter. Spread retries out so a struggling downstream system has time to recover, and so parallel workers don't all retry at the same moment:

function nextDelaySeconds(attempt):
    base = 2
    cap  = 900                                   // 15 minutes
    exp  = min(cap, base * 2^attempt)            // 2, 4, 8, 16 ... capped
    return random(0, exp)                        // "full jitter"

function process(message):
    try:
        handle(message)                          // must be idempotent
        ack(message)
    catch TransientError as e:
        if message.attempts < MAX_ATTEMPTS:      // e.g. 5-8
            requeue(message, delay = nextDelaySeconds(message.attempts))
        else:
            deadLetter(message, e, reason = "RETRIES_EXHAUSTED")
    catch PermanentError as e:
        deadLetter(message, e, reason = "NON_RETRYABLE")

In Apex, System.enqueueJob(job, delayInMinutes) lets a Queueable schedule its own retry with a delay of up to 10 minutes. For longer schedules, store the next retry time on the error record and use a scheduled job to pick it up.

Retries are only safe if processing is idempotent. Otherwise every retry is another chance of a duplicate. See designing idempotent ingestion layers.

Alerting and replay workflow

A DLQ that nobody looks at just loses data more slowly. Set up a workflow around it.

Alert on signals, not noise:

  • New permanent failures. Alert on any message dead-lettered for a non-retryable reason, grouped by error type so 300 identical validation errors arrive as one alert.
  • DLQ depth and age. Alert when the DLQ has more than a set number of messages, or when the oldest one is older than your agreed triage window.
  • Retry spikes. A sudden rise in retries often comes before an outage or a limit being exhausted.
  • Silence. No messages processed during a period when you'd normally see traffic usually means something upstream has broken.
  • Platform headroom. Daily API usage approaching its allocation, read from the Salesforce REST /limits resource.

Replay workflow:

  1. Triage. The on-call owner groups DLQ messages by error code and finds the root cause: a data issue, a mapping gap, a code defect or a configuration change.
  2. Fix the cause first. Correct the data, add the missing mapping or deploy the fix. Replaying before the cause is fixed just fills the DLQ again.
  3. Replay in a controlled way. Move messages back to the main queue (for example with SQS redrive) or re-run them with a replay script, starting with a small sample before the rest.
  4. Verify. Confirm the target records are correct in Salesforce or the downstream system.
  5. Close out. Mark each message as resolved, replayed or discarded, with who did it and why. Discarding should be deliberate and recorded.

Automatic or manual replay? Automatic redrive suits failures with a known temporary cause, such as replaying everything after a confirmed downstream outage ends. Manual or approved replay suits permanent errors, because someone has to fix the cause first. Many teams use both: automatic for "retries exhausted during an incident", manual for "non-retryable".

Logging fields to keep

Log enough to diagnose and replay without opening the raw payload, and store as little personal data as you can. That matters for UK GDPR and EU GDPR. For each message, consider keeping:

Field Purpose
correlation_id Traces one business transaction across every system and step
message_id / idempotency_key Makes replays safe and duplicates detectable
source_system, target_system Shows where it came from and where it was going
operation, entity_type For example "upsert", "Invoice"
external_id, salesforce_record_id Links to the business record in each system
attempt_count, first_attempt_at, last_attempt_at Retry history and age
error_category Transient or permanent, for routing and reporting
http_status, error_code, error_message The exact failure; truncate long messages
payload_ref or hash A pointer to the stored payload (encrypted, access-controlled) instead of the raw body in logs
schema_version Shows which mapping version processed it
status, resolved_by, resolved_at, resolution_note Audit trail for replay or discard

Set retention deliberately. Keep logs and DLQ payloads long enough for triage and audit, and no longer than your data protection policy allows.

FAQ

Frequently asked questions

How long should messages stay in a dead-letter queue?

Long enough to cover realistic detection and triage, including weekends and holidays, and no longer than your data retention policy permits. Platform limits apply too: Amazon SQS keeps messages for at most 14 days, for example. For longer audit needs, archive the error records to durable storage.

Should failed messages be replayed automatically or manually?

It depends on the failure. Automatic replay suits failures with a confirmed temporary cause, such as a downstream outage that has since ended. Permanent failures like validation errors or missing mappings need a person to fix the cause first, so replay those manually or with approval.

What should we alert on in an integration pipeline?

Alert on new non-retryable failures (grouped by error type), DLQ depth and age of the oldest message, unusual retry rates, unexpected silence in normally busy flows, and Salesforce API allocation nearing its limit. Route alerts to a named owner. An alert nobody owns tends to get ignored.

How many times should we retry before dead-lettering?

There's no universal number. Five to eight attempts with capped exponential backoff and jitter covers most short outages without delaying data for too long. Base it on how long your downstream systems usually take to recover and how much delay the business can accept.

Can Salesforce act as the dead-letter queue?

Yes, for Salesforce-centred designs. A custom error object holding the payload reference, error details and status can serve as a DLQ, with list views or reports for triage and a Queueable or scheduled job for replay. Watch your data storage, keep payloads small, and restrict access to sensitive fields.

Build recovery in from the start

The 5-day Integration Sprint (from €2,800) ships each integration with exponential backoff, a dead-letter queue and replay scripts, so temporary outages don't quietly lose records. For ongoing monitoring and triage of HTTP 429 and 5xx errors across your integrations, the Fractional Salesforce Architect retainer (€1,450/month) gives you a senior engineer who owns that work.