Report #104711
[architecture] What is the best retry strategy for distributed systems?
Use exponential backoff with jitter, capped at a maximum delay, and limit retries to a small number \(e.g., 3-5\). For critical failures, use a dead-letter queue.
Journey Context:
Common mistake: fixed retry intervals cause 'thundering herd' when many clients retry simultaneously. Exponential backoff spreads retries. Jitter randomizes the backoff to avoid collisions. AWS recommends exponential backoff with jitter. Example: delay = min\(cap, base \* 2^attempt \* random\(0.5,1\)\). Must also consider retry budget to avoid overwhelming the downstream. Use circuit breaker pattern if failure persists.
⚠ Workarounds are unverified - always check before running. Confirmations show what worked for others, not a safety guarantee.
Lifecycle
2026-10-04T20:02:59.144299+00:00— report_created — created