Why AI Workloads Are Forcing the Shift to Liquid Cooling
The Thermal Crisis Behind the AI Revolution
Artificial intelligence is reshaping every industry, but it is also reshaping something far less visible: the inside of your data center. Training large language models and running inference workloads at scale requires GPU clusters that generate heat densities that would have seemed extraordinary just five years ago. Today, a single GPU-dense rack can exceed 60 kW of heat load—a figure that exposes the fundamental limits of traditional air-based cooling and forces engineers, architects, and facility planners to rethink infrastructure from the slab up.
As your RCDD partner and a distributor focused on network infrastructure and data center products, Heather Technologies is seeing this shift play out in real procurement decisions every quarter. This article explains why liquid cooling is no longer a niche solution reserved for hyperscalers, and what the transition means for the broader data center ecosystem.
Why Air Cooling Hits a Wall
Conventional computer room air handlers (CRAHs) and precision air conditioning units are engineered around an assumption that traditional server racks rarely exceeded a few kilowatts. ASHRAE TC 9.9, the primary industry reference for data center thermal guidelines, establishes recommended IT equipment inlet temperatures in the range of approximately 18–27°C for Class A1 and A2 environments. Meeting those inlet conditions is straightforward at modest rack densities. At 60 kW per rack—or beyond—it becomes geometrically more difficult.
The core problem is the thermodynamic capacity of air itself. Air has a relatively low specific heat capacity compared to water. Moving enough conditioned air through a rack generating 60 kW of heat requires airflow volumes that create unacceptable noise, pressure differentials, and energy consumption. Hot spots emerge despite best-practice hot/cold aisle containment, and power usage effectiveness (PUE)—defined as total facility power divided by IT power—climbs as mechanical systems strain to compensate. Air cooling at extreme densities is not just less efficient; it eventually becomes physically inadequate.
Liquid Cooling: The Physics Work in Your Favor
Water and water-glycol mixtures carry heat away from processors orders of magnitude more effectively than air. This fundamental advantage makes liquid cooling the engineering answer to AI's thermal challenge. There are several deployment models, and they are often combined in hybrid configurations:
- Direct liquid cooling (DLC) / cold plates: Coolant circulates through cold plates mounted directly on CPUs and GPUs, capturing heat at the source before it ever becomes an airborne problem.
- Rear-door heat exchangers (RDHx): A liquid-cooled door mounted at the rear of a standard rack captures heat exhausted by servers. In passive configurations with EC fan assist, a single rack rear door can manage substantial heat loads, making RDHx a compelling retrofit path for existing facilities.
- Immersion cooling: Servers are submerged in dielectric fluid, enabling extremely high heat-flux dissipation. Single-phase and two-phase variants exist, each with different infrastructure implications.
- Coolant distribution units (CDUs): A CDU acts as the facility-side heat exchanger, transferring heat from the server-side loop (typically a propylene-glycol/water mixture) to the building chilled water or an external dry cooler loop.
In a representative 500 kW-IT containerized or modular edge AI data center, a CDU rated for approximately 350 kW can serve a cluster of high-density GPU racks, while rear-door heat exchangers handle supplemental loads. External dry coolers with adiabatic pre-cooling extend the hours of economizer-mode operation even in challenging ambient conditions, supporting a PUE target in the range of approximately 1.25—a meaningful improvement over air-only designs at equivalent densities.
Standards and Compliance Considerations
Liquid cooling introduces infrastructure complexity that must be addressed within the existing standards framework governing data centers.
Thermal and Site Design
ASHRAE TC 9.9 guidelines remain relevant even in liquid-cooled facilities because not every component in the rack—network switches, storage controllers, power supplies—is necessarily on the liquid loop. Hybrid designs must maintain compliant inlet air conditions for air-cooled components while routing liquid to GPU cold plates.
Infrastructure Ratings and Redundancy
ANSI/TIA-942 addresses data center infrastructure including power and cooling system design. Operators pursuing Uptime Institute Tier III certification—concurrently maintainable systems—must ensure that liquid cooling loops, CDUs, and dry coolers are designed with the same N+1 or 2N discipline applied to power. A CDU failure that takes down a GPU cluster is operationally indistinguishable from a UPS failure.
Electrical Infrastructure
High-density AI racks place extraordinary demands on power distribution. A 500 kW-IT facility may be served by a 480V three-phase system with an online double-conversion UPS configuration and intelligent rack PDUs providing per-outlet metering and dual A+B feeds. Automatic transfer switching integrating utility, solar, and battery energy storage adds resilience. NEC/NFPA 70 governs electrical installation and grounding throughout; Type 1 and Type 2 surge protective devices are required for transient voltage management. NFPA 70E arc-flash safety protocols must be followed during commissioning and maintenance of high-amperage distribution equipment. ANSI/TIA-607 addresses bonding and grounding infrastructure, which becomes especially important when introducing conductive fluid systems in close proximity to energized IT equipment.
Fire Detection and Suppression
Liquid cooling does not eliminate fire risk—it changes its character. Dielectric fluids have their own flammability profiles, and glycol-based systems introduce leak risk near electrical equipment. NFPA 75 covers protection of IT equipment, and NFPA 2001 governs clean-agent fire suppression systems. Facilities handling high-value AI workloads commonly deploy clean-agent suppression (such as Novec 1230 / FK-5-1-12) combined with VESDA aspirating smoke detection for the earliest possible warning.
The Business Case Is Already Here
Operators sometimes frame liquid cooling as a premium add-on. In the context of AI infrastructure, this framing is backwards. The cost of under-provisioned cooling—GPU throttling, shortened hardware lifespan, unplanned downtime, and stranded capacity—exceeds the capital investment in a properly engineered liquid cooling system. Energy savings from a materially lower PUE compound over the operational life of the facility.
Beyond economics, there is a market access argument. Cloud providers, enterprise AI teams, and government agencies specifying high-performance computing infrastructure increasingly mandate or prefer liquid-cooling-capable facilities. A data center that cannot support liquid cooling is progressively unable to serve the highest-value workloads in the market.
What This Means for Your Next Project
Whether you are designing a new edge AI facility, retrofitting an existing data center, or specifying equipment for a colocation build-out, the engineering conversation must start with heat load per rack—not total facility power. At densities above roughly 20–25 kW per rack, hybrid liquid cooling deserves serious evaluation. Above 40 kW per rack, it is functionally mandatory.
Heather Technologies works with engineers and data center operators to source the CDUs, rear-door heat exchangers, intelligent PDUs, containment systems, and supporting infrastructure needed to build AI-ready facilities that meet both today's workloads and the denser, hotter hardware on the near-term roadmap. The thermal physics of AI are not a trend—they are a permanent shift in what infrastructure must deliver.
About the author — Todd Taskerud, AWS CCP, RCDD/NTS/OSP/WD, LEED GA, is a BICSI-credentialed communications distribution designer at Heather Technologies, specializing in fiber, copper, and data-center network infrastructure.