Getty Images/iStockphoto

AI data centers can run hot to stay cool: report

Schneider Electric reports it found room to run AI data centers at a higher temperature and lower direct on-site water use, but there are trade-offs.

AI data center operators can cut direct onsite water use by sending warmer water through their liquid-cooling systems, according to a Schneider Electric analysis of 100-megawatt facilities.

Raising technology cooling system (TCS) supply temperature from 32 C (89.6 F) to 45 C (113 F) lowered its water-usage effectiveness (WUE) ratio by about 66% in Dallas and 82% in Paris, Schneider found. The comparison was with its 32 C liquid-cooling design.

The warmer loop gives dry coolers a larger temperature difference to reject heat. That lets the system rely less on water-consuming evaporative cooling and delays the need for chiller assistance.

The tradeoff is equipment. The water-saving design can require more dry-cooler capacity and higher capital spending. Operators also must confirm that their server, rack and liquid-cooling configurations are approved for the relevant higher inlet-water temperatures.

The 45 C figure refers to water supplied through the cooling system, not the temperature of the GPUs themselves. The water reductions were far larger than the associated power usage effectiveness (PUE) gains, which were below 1% in Schneider's 45 C comparisons.

"From a water use and efficiency perspective, high temperature liquid cooling makes a lot of sense," said Steve Carlini, chief advocate, AI and data centers at Schneider Electric.

Schneider modeled a 100 MW AI factory using Nvidia Vera Rubin liquid-cooled clusters. It compared a traditional air-cooled design with three liquid-cooling configurations, including systems operating at 32 C and 45 C TCS supply temperatures.

Schneider compared four cooling configurations: a traditional air-cooled design, a 32 C liquid-cooling design, a 45 C liquid-cooling design with the same cooling equipment as the 32 C system, and a 45 C design with fewer cooling units. The scenarios let Schneider compare water use, energy efficiency and equipment requirements as operators raise the TCS supply temperature.

Higher rack densities raise the stakes

Schneider's AI factory scenarios modeled rack densities of up to 188 kW, compared with 20 kW and 40 kW in its traditional air-cooled scenario. The AI designs used separate liquid- and air-cooling loops, with about 86.6% of the IT load liquid cooled.

As AI rack densities rise, warmer facility-water loops give operators another way to limit direct onsite water use while supporting high-density liquid cooling.

For AI campuses facing water constraints, 45 C is a design option that can trade additional dry-cooler capacity and capital spending for lower direct water use. Climate, water availability, cooling architecture, hardware ratings and operator risk tolerance will determine whether that trade works for a particular facility.

Warmer water extends dry cooling

Water savings come from using dry coolers more often and relying less on evaporative cooling.

Paris's lower peak temperatures give the 45 C design more hours to reject heat with dry coolers without adiabatic assistance. Schneider modeled outdoor temperatures ranging from minus 7 C to 29 C in Paris, compared with minus 9 C to 39 C in Dallas.

In Schneider's optimized 45 C design, chiller assistance starts at about 35 C outdoors, compared with about 25 C at 32 C.

Adiabatic cooling uses evaporating water to improve heat rejection when outdoor temperatures rise.

"The higher TCS temperature shifts to more time operating on the dry coolers in Paris versus Dallas without adiabatic assist, thus using less water," Carlini said.

Moving from traditional air cooling to 32 C liquid cooling cut onsite water use in Dallas by about 8% while improving PUE by roughly 7%.

Less water can mean more cooling equipment

Schneider's third scenario kept the same cooling equipment used in its 32 C design but raised the TCS temperature to 45 C.At 45 C, about 33% of the installed cooling capacity was surplus to the modeled requirement. Schneider treated that additional dry-cooler surface area as intentional overprovisioning designed to reduce reliance on water-consuming cooling.

"While it may be called 'stranded' it's a preference to have as much surface area as possible for the dry coolers to operate and minimize water use," Carlini said. "When water use is the main focus, this CapEx is justified."

It's a preference to have as much surface area as possible for the dry coolers to operate and minimize water use. When water use is the main focus, this CapEx is justified.
Steve Carlini, Chief advocate, AI and data centers, Schneider Electric

Schneider also modeled a fourth scenario that reduced the number of cooling units while keeping the 45 C supply temperature. Compared with the 32 C liquid-cooling design, that configuration cut WUE by about 47% in Dallas and 53% in Paris.

It also reduced capital costs while using more water than the fully equipped 45 C design, Carlini said.

Schneider also modeled a fourth scenario that reduced the number of cooling units while keeping the 45 C supply temperature. Compared with the 32 C liquid-cooling design, that configuration cut WUE by about 47% in Dallas and 53% in Paris, substantially less than the fully equipped 45 C design.

In other words, operators can reduce the amount of cooling equipment they install and lower capital costs, but they give up some of the water savings that come from overprovisioning dry-cooler capacity.

The scenarios show that operators can tune the design for capital cost, water consumption, operating costs or redundancy -- but cannot maximize all four at once.

Hardware limits still apply

A 45 C supply temperature expands the range of options for liquid-cooled AI facilities. It does not make 45 C a universal operating target.

CoolIT Systems said 45 C is the maximum TCS temperature approved for Nvidia hardware using warm-water cooling.

"It is not necessarily that 45 C exactly is the target, but it opens up the allowable envelope to a higher maximum temperature compared to what people were commonly looking at previously," said Charles Robison, CoolIT's director of marketing, relaying comments from the company's engineering team.

Thomas Loxley, senior manager of codes and standards at ASHRAE, said 45 C supply water can provide enough temperature difference to cool IT equipment, depending on the liquid-cooling design.

Higher temperatures can affect hardware component life, however.

"Silicon and other power distribution components can experience accelerated aging at higher temperatures," Loxley said. Newer equipment rated for higher temperature classes should be considered when changing operating temperatures, he said.

Carlini said Nvidia supports and warrants its next-generation architecture, including the Vera Rubin platform and compatible MGX architectures, for operation at a 45 C (113 F) facility/rack inlet water temperature.

For other vendors, selected models and server configurations support continuous or temporary ASHRAE A4 operation, which permits inlet temperatures up to 45 C, he said.

Even when hardware supports the temperature, operators may choose a lower setpoint.

 Carlini said data center operators have traditionally run GPUs at lower temperatures, and some may continue to do so because of service-level agreements, risk tolerance and concerns about long-term reliability.

Direct water isn't the whole picture

In the Paris scenarios, adding adiabatic assist improved PUE by an average of just 0.13%, while increasing water consumption from zero to 0.07 liters per kilowatt-hour.

Schneider also distinguishes direct on-site cooling water consumption from water used to generate electricity. That indirect water use depends on the local power mix and was outside the scope of the cooling simulations.

Closed-loop liquid cooling can allow facilities to operate with little or no on-site cooling water consumption. Coolant may require glycol to prevent freezing, Carlini said, but the water can remain in a closed system for decades before replacement.

Shane Snider is a senior news writer at TechTarget, covering AI infrastructure, hyperscale data centers, cloud platforms, and the power and energy systems driving modern compute expansion. You can reach Shane at [email protected] or on LinkedIn.