LAB-001DISTRIBUTED SYSTEMS

What changes when an automated action fails halfway through?

PUBLISHED

Most automation is written to describe the happy path: do step one, then step two, then step three. Failure handling, when it exists at all, usually means “stop and alert someone.” That’s adequate when a human is the next line of defense. It’s not adequate when the system is expected to recover on its own.

The interesting failures aren’t the clean ones. A process that fails before touching anything is easy to reason about — just retry it. A process that fails after every step completes is also easy — it already succeeded. The hard case is the one in the middle: three of five steps committed, the fourth failed, and the system now has to decide what “correct” even means for the two steps that never ran.

We spent time cataloguing where this actually bites, rather than theorizing about it. The pattern that showed up repeatedly: operations that look atomic from the outside are usually a sequence of smaller operations underneath, each with its own failure window. Infrastructure provisioning is a good example — allocating a resource, configuring it, registering it, and exposing it are four separate steps that can each fail independently, and most tooling doesn’t model that the attempt partially succeeded until asked to clean up.

Two approaches showed up in how existing systems handle this, and neither is a complete answer on its own:

  • Compensating actions — for every forward step, define the action that undoes it, and run those in reverse on failure. This works cleanly when every step has an obvious inverse. It works badly when a step has side effects that can’t be fully undone — a notification already sent, a charge already processed.
  • Idempotent resumption — design each step so that re-running it from the top is safe, and let the system simply retry the whole sequence until it completes. This avoids needing an inverse for every action, but it requires every step to tolerate being run more than once, which is a real constraint on how each step is written.

Neither approach is free. Compensating actions require designing an inverse for every forward action before you need it. Idempotent resumption requires discipline at every step, forever, since one non-idempotent step anywhere in the chain breaks the guarantee for the whole sequence.

What we didn’t find: a shortcut. The systems that handle partial failure well treat it as a design constraint from the first step, not a bug to patch in after the fact.


← All lab notes