Back to Insights
Jamal Osei

AI Inference Clusters and the 24/7 Power Requirement

GPU server rack interior with glowing indicator lights

The phrase "machine learning workload" covers a lot of ground. For years, the dominant framing was training: large-scale compute clusters running for weeks or months on curated datasets, iterating toward a model. Training runs are expensive and time-consuming, but they have a property that makes power provisioning relatively forgiving. They are batch jobs. If your cluster loses power at 3 AM on a Tuesday, you restart from the last checkpoint. You lose compute hours. That is a real cost, but it is not a service failure. No customer sees an error.

Inference does not work that way.

When a model is deployed in a production application, whether that is a search product, a code assistant, an autonomous agent workflow, or a real-time document processing pipeline, the compute has to be available when the request arrives. There is no checkpoint to fall back to. A user made a request, an API call fired, and within tens to hundreds of milliseconds your cluster needs to return a result. If the power is unstable and your GPU nodes are throttling or offline, the request times out. That is not a power event. It is a service failure, and it shows up directly in availability metrics.

The Training-Inference Power Profile Divergence

The power demand characteristics of a busy inference cluster look different from a training cluster, and that difference matters when thinking about what kind of power source is appropriate. Training workloads ramp up, hit sustained high utilization for a defined period, then drop. You can, to some extent, schedule training jobs around expected power availability windows. If you know grid capacity is constrained between 4 PM and 8 PM on hot summer days in your region, you schedule the next training run to start at 9 PM.

Inference clusters at scale have a near-flat, continuously high load profile. The GPUs are occupied. The interconnect fabric is running. Cooling systems are at or near design load. For a facility serving live products with real user traffic, this load does not disappear at night or on weekends. A financial services firm running real-time fraud detection on transaction streams sees steady load at 2 AM on Saturday because transactions happen continuously. A data platform running continuous inference against incoming documents has no off-peak window.

This is the fundamental shift: AI products moved from research pipelines to operational infrastructure, and the power requirements moved with them.

What 99.99% Uptime Actually Requires

Infrastructure teams managing inference clusters work with uptime targets, not power targets. Those targets translate into very small annual outage budgets. A 99.9% uptime specification allows roughly nine hours of downtime per year. A 99.99% target allows about 52 minutes across the full year. When teams are operating production AI products with SLAs in that range, every unplanned power event counts against the budget.

Grid-supplied power in most regions is not rated to those reliability levels. Transmission-level availability statistics from utility reporting cover full outages measured in hours, not the brownouts, voltage sags, and frequency excursions that affect compute performance even without a complete outage. A data center that relies entirely on grid power and battery UPS for ride-through has a reliability ceiling that is hard to raise further with conventional approaches.

Battery backup systems handle short-duration outages well. A well-designed UPS installation can carry a facility through a two-to-four-minute grid disturbance while a backup generator spins up. But battery storage at the scale required to cover a multi-hour grid curtailment for a facility drawing 50 megawatts or more is not practically deployable today, either economically or physically. The battery installation would dwarf the compute facility itself.

Power Quality, Not Just Power Availability

There is an effect of grid instability that is less visible than a full outage but affects inference performance in ways that are hard to debug: power quality. Modern GPU compute clusters are sensitive to voltage and frequency stability, not only to whether power is present at all. When input voltage drops below design thresholds, power supply units derate their output. GPU clock speeds and memory bandwidth drop. Inference latency increases. For latency-sensitive applications, this degradation surfaces as intermittent tail latency spikes that are difficult to attribute because they are not caused by software bugs or network issues.

We look at this problem closely because it affects the specification we are setting for our own power delivery architecture. A fission plant with well-controlled load-following behavior delivers stable voltage and frequency to the data center bus it connects to. The compute cluster does not see the variability that propagates from upstream grid conditions. Latency performance becomes predictable in a way that matters for SLA management and capacity planning.

A Concrete Scenario

Consider a 60 MW inference facility operating in a hot-climate region. In summer, the local utility may issue demand-response events requesting that large industrial customers curtail load during peak afternoon hours. A data center with a grid-only supply either curtails compute capacity during those windows, which means dropping inference jobs or rejecting API requests, or it pre-qualifies for emergency reliability contracts, which are expensive and not universally available. A facility with dedicated on-site generation has no demand-response obligation to the grid and maintains full compute availability through the event.

We are not claiming every inference facility faces this exact scenario. Utilities and grid operators have different demand-response programs with different participation rules, and many large data center operators have negotiated contracts that exempt compute infrastructure from curtailment requests. But the underlying constraint is real: grid-dependent power ties inference availability to decisions made by parties whose interests are not aligned with yours.

Rethinking the Power Procurement Frame

For data center operators planning new inference capacity, the useful question is not "how many megawatts can we procure at what price?" The useful question is "how many of those megawatts will be physically available at any given moment, with what level of confidence, and what is the cost when they are not?"

Grid power procurement answers the first question and gives probabilistic answers to the second and third. A dedicated on-site generation asset designed as the primary power source for a specific facility answers all three with different levels of certainty. The asset either runs or it does not, the operational state is directly observable, and the availability is not contingent on grid conditions upstream.

This is the reasoning that shapes our work at Applied Atomics. We are not building a supplement to grid power. We are building a primary power source for facilities where inference SLAs are tight, load profiles are continuous, and the reliability ceiling of grid-dependent power is a real operational constraint. The conversation about whether on-site fission is the right choice for a given facility starts with the power requirement, not the power supply technology.

The Honest Timeline Constraint

Inference facilities being commissioned today will rely on grid power and whatever backup systems they can deploy. That is the practical reality, and we are not pretending otherwise. On-site fission takes years from early conversations to operational power. Operators who are standing up inference capacity in 2026 and 2027 are not going to be powered by compact fission units on that timeline.

What we are focused on is the planning horizon for facilities that are being designed now for operation in 2030 and beyond, where the scale of AI compute is large enough to justify dedicated power infrastructure, and where the reliability requirements are tight enough that grid dependence carries a real business risk. That is a different conversation, and it requires starting several years before first power is needed.