The Signal

For most of computing history, “where does this run” had a short answer: on the server you provisioned for it. Capacity planning meant buying or renting more of the same kind of machine.

That answer no longer holds. Organizations now operate across on-premises hardware, multiple cloud regions, specialized accelerators, and edge locations — often for the same application. The question has changed from where the server is to where a given workload should run, and that question can have a different answer every time it’s asked.

The Problem

A fixed deployment model optimizes for one thing: simplicity. It does not optimize for cost, latency, data locality, or availability, because it never asks the question. When compute is treated as a placement decision instead of a fixed location, the infrastructure has to weigh real constraints against each other every time it schedules work:

  • Latency — how close does this need to be to the user or the data?
  • Cost — what does this specific unit of compute cost in this specific location right now?
  • Memory — does the workload fit the available memory profile without expensive swapping or resizing?
  • Data locality — does moving the computation cost less than moving the data?
  • Privacy — are there constraints on where this data or computation is allowed to physically reside?
  • Accelerator availability — is specialized hardware available where and when it’s needed?
  • Power — are there limits on energy draw at a given location or time?
  • Policy — are there organizational or regulatory rules that constrain placement regardless of cost or performance?

The System Question

Placement decisions require infrastructure that can observe these constraints continuously and route accordingly, rather than infrastructure that was configured once and left alone. That is a scheduling and routing problem, not a provisioning problem, and it changes what “infrastructure” needs to know about a workload before it runs.

The Tradeoffs

Dynamic placement adds real engineering complexity: more moving parts, more failure modes, more to monitor. A workload that doesn’t vary — a service with stable, predictable load and no locality or privacy constraints — may not benefit enough to justify that complexity. The decision to treat compute as placement is itself a tradeoff, not a default.

Business Impact

Treating compute as a placement decision allows organizations to optimize infrastructure around economics and requirements rather than around a single deployment model.

The strategic advantage is not owning every resource. It is having the ability to choose the right resource at the right time.

What We’re Watching

Heterogeneity is the leading indicator: the more types of compute an organization operates — different accelerator generations, different regions, different ownership models — the more a fixed deployment model leaves on the table. Watch for organizations running mixed fleets as the population most likely to need placement-aware infrastructure first.

The infrastructure question is becoming less about where compute exists and more about where computation belongs.