Skip to main content

Consulting Specifying Engineer January February 2026

Page 16

BUILDING SOLUTIONS EMERGENCY, STANDBY, BACKUP

Sam Buscemi, PE, NCEES, Affiliated Engineers Inc., Madison, Wisconsin

How to achieve AI data center backup power, cooling

AI data centers have redefined the power paradigm in design, shifting from ensuring uninterrupted uptime to prioritizing power stability and cooling continuity.

F

Learning

Objectives

u

•D ifferentiate between backup power expectations in traditional colocation data centers versus AI data centers, including how reliability targets and risk profiles shape system design. •U nderstand how design priorities in AI data centers, including maintaining cooling loops, protecting valuable hardware and enabling controlled shutdowns, are changing the application of uninterruptible power supply, battery energy storage and hybrid backup systems. •E valuate alternative power design strategies that balance resiliency, emissions/permitting hurdles and practical equipment availability for supporting AI workloads.

or decades, the benchmark of excellence in data center design was simple to state and expensive to achieve: keep everything on, all the time. Extensive fleets of diesel generators, redundant uninterruptible power supply (UPS) systems and mechanical plants enabled data centers to operate indefinitely without the grid. Downtime for financial institutions, hospitals or e-commerce platforms translated into revenue loss, compliance risk or threats to public safety. What mattered most was keeping services running. Hardware was expendable, but uptime was not. The rise of artificial intelligence (AI) has shifted the equation. AI data centers concentrate dense racks of graphical processing units (GPUs) that can draw up to 130 kilowatts each, cooled by direct-tochip liquid systems with tight flow and temperature tolerances. A single high-density AI compute cabinet is estimated to cost between $2 million and $3 million and when multiplied across the hundreds of units deployed in a modern facility, the hardware investment alone quickly climbs into the hundreds of millions. Protecting that capital investment is no longer just an operational concern — it is a business imperative.

Uptime versus asset protection AI facilities operate under fundamentally different economics than traditional colocation providers. Training workloads can be checkpointed and resumed, making brief interruptions tolerable. Customers expect occasional capacity constraints, such

14 | January/February 2026 CSE2601_MAG_MCF_V3msFINAL.indd 14

as when new model launches prompt providers to throttle resources to manage demand. Some downtime is acceptable, but hardware damage is not. The principal responsibility for consulting engineers in AI data centers has shifted from maximizing uptime to protecting GPU assets and preventing catastrophic loss. The reason is twofold. First, each GPU represents an enormous upfront investment and once damaged, the capital loss is permanent with no recovery path. Second, the direct-to-chip liquid cooling used in these systems introduces an acute vulnerability: without continuous coolant flow, the risk of thermal runaway is immediate and severe. This stands in sharp contrast to the traditional air-cooled data hall. In those environments, a loss of cooling might mean a gradual rise in data hall temperature over minutes or even hours, allowing workloads to be throttled down or shifted before equipment becomes at risk. With GPUs, however, sudden power interruptions or cooling failures can push hardware past safe limits in seconds, causing thermal overload, electrical stress and cascading failures across densely packed racks. As a result, the focus has shifted away from uptime at all costs toward ensuring that systems remain stable long enough for workloads to throttle down and shut off safely. This evolution forces a reconsideration of what “mission critical” entails, with resiliency measured by the ability to protect GPUs rather than keep every workload running indefinitely. Instead of covering every kilowatt of information technology (IT) and cooling load with diesel generator backup, the industry is shifting toward hybrid strategies, including battery energy storage systems (BESS), thermal energy storage, selective generator coverage and advanced protection schemes and smoothing load fluctuations of GPU servers. These designs now prioritize protecting high-value hardware investconsulting-Specifying engineer — www.csemag.com

12/23/25 10:50 AM


Turn static files into dynamic content formats.

Create a flipbook
Consulting Specifying Engineer January February 2026 by Arrowfly - Issuu