Retry with Exponential Backoff & Jitter
Key Learning Objectives
Implement production-grade asynchronous retry loops with bounded maximum attempts.
Formulate exponential backoff intervals with full jitter to mitigate thundering herd spikes.
Differentiate retryable server faults (5xx, network timeouts) from non-retryable client errors (4xx).
Integrate AbortSignal and request cancellation to immediately halt backoff delays on user abort.
Safeguard non-idempotent operations against duplicate execution hazards using idempotency keys.
The Interview Problem
What is logged to the console when the following retry utility executes against an API that fails twice with 503 errors before succeeding, using an injected deterministic sleep provider?
1async function retryWithBackoff(fn, options = {}) {2 const { maxRetries = 3, baseDelay = 100, shouldRetry = () => true, sleep = (ms) => new Promise(r => setTimeout(r, ms)) } = options;3 let attempt = 0;4 while (true) {5 try {6 return await fn(attempt);7 } catch (err) {8 attempt++;9 if (attempt > maxRetries || !shouldRetry(err)) {10 throw err;11 }12 const delay = baseDelay * (2 ** (attempt - 1));13 await sleep(delay);14 }15 }16}1718const delays = [];19const mockSleep = (ms) => { delays.push(ms); return Promise.resolve(); };2021let callCount = 0;22const flakeyApi = async (attempt) => {23 callCount++;24 if (callCount < 3) {25 const err = new Error('503 Service Unavailable');26 err.status = 503;27 throw err;28 }29 return 'SUCCESS';30};3132const isRetryable = (err) => err.status >= 500;3334retryWithBackoff(flakeyApi, { maxRetries: 3, baseDelay: 50, shouldRetry: isRetryable, sleep: mockSleep })35 .then((result) => {36 console.log(result, callCount, delays.join(','));37 });
Predict Console Output
Select the option that matches what standard ECMAScript prints to the console:
SUCCESS 3 50,100
SUCCESS 3 100,200
SUCCESS 2 50
Error: 503 Service Unavailable 3 50,100
V8 Engine Execution Trace
Step 1 of 8 (Line 1)Defines retryWithBackoff and test fixtures in global environment.
Deep Technical Breakdown
Resilient Asynchronous Retries in Production
Blindly retrying failed network requests is one of the quickest ways to cause cascading backend outages (the Thundering Herd problem). A production-grade retry implementation requires four foundational pillars:
1. Exponential Backoff Formula
Increasing wait times exponentially prevents overwhelming a recovering server:
$$\text{delay} = \min(\text{maxDelay}, \text{baseDelay} \times 2^{\text{attempt}-1})$$
2. Full Jitter (Desynchronization)
When thousands of clients fail simultaneously, pure exponential backoff causes all clients to retry in synchronized waves. Adding Full Jitter randomizes delays across the backoff window:
const delay = Math.random() * (baseDelay * (2 ** (attempt - 1)));AWS research shows full jitter delivers the lowest overall completion times and zero retry synchronization spikes.
3. Classification of Errors (Retryable vs Non-Retryable)
- NEVER Retry: 4xx client errors (400 Bad Request, 401 Unauthorized, 403 Forbidden, 404 Not Found, 422 Unprocessable Entity). Retrying a bad syntax request will never succeed and wastes battery/bandwidth.
- SAFELY Retry: 5xx server errors (500, 502 Bad Gateway, 503 Service Unavailable, 504 Gateway Timeout), transient network drops (TypeError: Failed to fetch), and Rate Limits (429 Too Many Requests, respecting
Retry-Afterheaders).
4. Idempotency & AbortSignal Integration
- Idempotency: Non-idempotent mutations (e.g.,
POST /payments) must not be retried without anIdempotency-Keyheader, or a network timeout could result in duplicate financial charges. - Cancellation: If a user navigates away or cancels the operation, the retry sleep should reject immediately via
signal.addEventListener('abort', ...)rather than hanging until the timer completes.
Common Traps & Mistakes
Retrying 4xx client errors such as 400 or 404, which can never succeed without request payload modifications.
Omitting jitter, causing all disconnected clients to flood the server simultaneously in synchronized thundering herds.
Blindly retrying non-idempotent POST operations without idempotency tokens, creating duplicate records or payments.
Failing to clean up sleep timers when an AbortSignal aborts, leading to unhandled promise rejections and memory leaks.
FAANG Follow-Up Probes
Probe #1
How would you integrate standard HTTP 'Retry-After' response headers into the backoff delay calculator?
Probe #2
How does Decorrelated Jitter differ from Full Jitter, and when would you choose one over the other?
Probe #3
How would you write a circuit breaker pattern wrapper around this retry utility to fail fast during prolonged outages?
