June 15, 2026 · 6 min read
Your liquid cooling loop is a flow assurance problem
Facilities teams are treating coolant networks as plumbing. The physics says otherwise.
Direct-to-chip liquid cooling arrived in AI data centers faster than the discipline needed to design it safely. Racks now exceed 100 kW, and the coolant network keeping them alive is still specified with steady-state spreadsheets: a pump curve here, a pressure drop calculation there, a safety factor tacked on at the end.
That is the same starting point the oil and gas industry had with multiphase pipelines thirty years ago, before flow assurance existed as a discipline. It took a string of transient failures, pigging operations gone wrong, slug flow nobody modeled, waterhammer that split pipe, before steady-state design was recognized as necessary but not sufficient.
Liquid cooling networks for AI data centers are at that same inflection point.
Where steady-state analysis stops seeing the problem
A pump curve tells you what the system does at equilibrium. It says nothing about what happens in the seconds after a pump trips, a valve slams shut, or a rack load steps from idle to full draw. Those are exactly the events that matter, because they are the ones that damage hardware, trip protection systems, or take a hall offline.
Three specific blind spots recur across the loops we have reviewed:
- Maldistribution. Vendor curves assume even flow split across parallel branches. Real manifolds do not split evenly, especially under partial load or after a valve reconfiguration, and the branches starved of flow are exactly the ones running the hottest chips.
- Two-phase instability. As operators push toward higher exit temperatures and lower flow rates to save pumping energy, coolant moves closer to its saturation point. Once local boiling starts, the flow regime can become unstable in ways a single-phase steady-state model cannot represent.
- Waterhammer. A CDU pump trip or a fast-closing isolation valve generates a pressure transient that propagates through the network in milliseconds. Depending on pipe run length and fitting layout, that transient can exceed design pressure well before any operator or control system reacts.
What transient analysis actually requires
None of this is new physics. It is the same transient thermal-hydraulic problem that pipeline engineers have solved for multiphase gathering systems for decades: mass, momentum and energy conservation solved on a network of nodes, with pump trip curves, valve closure profiles, and check valve behavior modeled explicitly rather than assumed away.
What is new is applying that rigor to a coolant network with hundreds of cold plates, dozens of CDUs, and failure modes measured in seconds rather than minutes.
That is the gap ignzai exists to close: network-level transient simulation, not room-level CFD, built for the specific topology and specific failure modes of liquid cooling in AI data centers.