LAB-002DISTRIBUTED SYSTEMS

Why retrying infrastructure operations is harder than it looks

PUBLISHED

“Just retry it” is the first instinct for most transient failures, and it’s usually the right one for read operations. It gets considerably more complicated for anything that changes state.

The obvious risk is duplication: retry a request that actually succeeded but whose response got lost, and you might perform the same action twice. The less obvious risk is timing — a retried operation can race against the original attempt if the original wasn’t actually dead, just slow. Both problems have the same root cause: the caller doesn’t actually know what happened on the other end, only that it didn’t get a confirmation in time.

A few things we found worth separating explicitly, rather than treating “retry” as one concept:

  • Request identity. Did this retry carry the same idempotency key as the original attempt, so the receiving system can recognize it as a duplicate rather than a new request?
  • Effect visibility. Can the caller check whether the original attempt’s effect already happened, before deciding to retry at all?
  • Backoff behavior. Is the retry spaced out enough to avoid making a struggling system worse, especially when many callers are retrying at once?
  • Retry budget. Is there a limit on how many times, and for how long, retrying is even the right response — as opposed to surfacing the failure?

None of these are exotic. What’s easy to miss is that they all depend on the receiving system being designed to answer them. A retry strategy is only as good as the receiving infrastructure’s ability to say “I’ve already done this” or “that’s still in progress” — without that, retrying is a guess, not a recovery mechanism.


← All lab notes