Hard20 minDistributed Systems
UpdatedAug 6, 2026
Edit

gRPC Deadlines and Retries

Question Variations

  • "Why should every gRPC call have a deadline?"
  • "Which gRPC status codes are candidates for retry?"
  • "How can retries cause an outage to worsen?"
  • "How would you make a reserve-inventory RPC safe to retry?"

Why This Is Asked

This tests whether a candidate can make RPC calls fail fast and recover safely under partial failure. Interviewers expect deadline propagation, idempotency-aware retries, backoff, and an understanding of retry amplification.

Key Concepts

  • Deadlines: Bound the full request budget and propagate it to downstream calls.
  • Retry eligibility: Retry only transient failures and operations whose effects can safely be repeated.
  • Backoff and jitter: Spread retries to avoid synchronized load spikes.
  • Failure containment: Combine timeouts, circuit breaking, and load shedding to prevent cascading failures.

Question Variations

  • “Why should every gRPC call have a deadline?”
  • “Which gRPC status codes are candidates for retry?”
  • “How can retries cause an outage to worsen?”
  • “How would you make a reserve-inventory RPC safe to retry?”

Answers by Technology

+ Add Variant
System DesignImprove this answer ✏️

Expected Answer

Every gRPC call should have a deadline that represents the remaining end-to-end latency budget. The receiving server exposes cancellation when the deadline expires; it must stop work and propagate the remaining budget to any downstream RPCs. Resetting a 500 ms timeout at every hop permits a short request path to run for seconds and exhausts resources during a dependency slowdown.

Retry only a bounded set of transient errors, such as an unavailable endpoint, and only when the request is idempotent or protected by an idempotency key. Use exponential backoff with jitter and a retry budget. Do not retry client errors, expired deadlines, or operations whose effects cannot safely happen twice. Combine retries with circuit breakers, concurrency limits, and load shedding: during an outage, naïve retries multiply traffic precisely when the dependency has least capacity. Observe attempts, status codes, and deadline expirations separately from final failures so policy can be tuned from evidence.

Why It Matters

Deadline-free calls and unbounded retries turn one slow service into a cascading outage. Correct policy protects callers and gives degraded dependencies time to recover.

Example Code

const deadline = new Date(Date.now() + 300);
try {
  return await inventory.reserve(request, { deadline });
} catch (error) {
  if (error.code === "UNAVAILABLE" && request.idempotencyKey) return retryWithJitter(request, deadline);
  throw error;
}

Common Mistakes

  • Using an unlimited client default: Hung calls retain sockets and work long after the user has left.
  • Retrying beyond the original deadline: The request is no longer valuable and only adds load to a failing dependency.

Follow-up Questions

  • What does deadline propagation prevent? (Answer: Each hop consuming a full timeout independently and exceeding the caller’s total latency budget.)
  • Why add jitter? (Answer: It prevents many clients from retrying in synchronized bursts.)

Related Questions

References