← All articles

Same Power, More Compute

How Fluidstack’s Demand Management system lets us deploy more compute with the same power.

This is the first post in a series on Fluidstack's demand management system. It covers what power oversubscription is and why it matters.

In order to deploy more compute, you typically need more power. In the world today, power is getting increasingly hard to come by. So how can we deploy more compute with the power that we already have? This is the question Fluidstack's Demand Management system seeks to answer.

What is power oversubscription?

Power oversubscription refers to installing hardware whose combined peak power draw exceeds the available capacity.

Consider a site with 30 megawatts available to power IT equipment. Conventional design would cap installed peak capacity at 30 megawatts, as it assumes that every accelerator could draw its peak power at once.

Let’s say instead that we were to deploy 40 megawatts of peak IT capacity. If all IT equipment were to peak simultaneously, power consumption would exceed the available 30 MW capacity, tripping protective breakers.

In practice, however, different AI workloads have different power profiles. That means that the power consumption of a cluster running a mix of workloads will fluctuate as workloads move through different phases – for example, inference power consumption rises and falls with traffic, while training power consumption varies depending on phase.

The facility must support the fleet’s actual demand at every moment, which is typically lower than the sum of every device’s peak power draw. When these peaks do coincide, the system must be able to throttle workload power consumption to ensure the facility’s limits are not breached.

Chart of site power over time. Two workloads peak at different times and stay below a 30 MW facility limit; when their peaks align, uncontrolled demand would cross the limit, but Demand Management intervenes and holds actual site demand just below it
Demand Management holds aggregate demand below the facility limit as workload peaks align.

The importance of reliability and controllability

Oversubscription means the site now operates closer to the facility’s protective limits. If telemetry is wrong or control reacts too slowly, protective breakers may trip and take equipment offline – the system must act before protection does. Fluidstack’s Demand Management system ensures this by continuously measuring facility telemetry, and actively controlling demand when thresholds are about to be breached.

The facility needs an accurate and reliable understanding of current power and cooling limits that its active equipment can support. These limits can change – for example with transfers to backup generation, failures in cooling systems, or maintenance operation. The system must compare aggregate demand against the current boundary and act before it crosses it. If telemetry or control becomes unavailable, the system must also fall back to a known safe operating limit.

The solution crosses the facility boundary

Demand Management depends on:

  • Electrical and cooling infrastructure adequately sized for the intended deployment.

  • Fresh, reliable, and granular telemetry on current power and thermal conditions, equipment state, and alarms.

  • An accurate live topology model which connects each load to the upstream power and cooling infrastructure which serves it, and how these domains overlap.

Together, these pillars help us to observe the facility, identify active constraints, calculate new safe limits, communicate them to the scheduler, and verify that demand changed.

Making intelligence abundant involves both adding new power generation, and also making better use of the existing capacity we have. Demand Management allows Fluidstack to deploy a larger fleet by using otherwise unused capacity.

The next two posts explain the underlying telemetry and topology systems that enable this.

Aaron Jeyaraj
Software Engineer