July 2, 2026 · 5 min read
What happens in the 10 seconds after a CDU pump trips
A pump trip is not a steady-state event. Here is the transient sequence most designs never check.
A coolant distribution unit pump trip is one of the most common transient events a liquid-cooled data hall will see, and one of the least tested at design time. Most designs verify that N+1 pump redundancy exists. Few verify what the network does in the seconds between the trip and the standby pump reaching full flow.
The sequence
T+0 seconds. The duty pump trips, on a fault, a power blip, or a manual isolation. Flow through the CDU begins to fall immediately, but not uniformly across the network: branches closest to the pump lose head first, while distal branches coast briefly on stored momentum.
T+0 to T+2 seconds. Static pressure at the pump discharge collapses. If a check valve slams shut against reverse flow, the resulting pressure transient can exceed the steady-state design pressure at fittings and manifold junctions well before this window closes. This is the waterhammer spike that steady-state design never sees, because it does not exist in a steady-state model.
T+2 to T+6 seconds. Flow redistributes across the remaining active paths. Branches with the least resistance take a disproportionate share of whatever flow remains, which is rarely the branch carrying the hottest rack. Local flow through the highest-power cold plates can drop faster than average network flow, because maldistribution under transient conditions does not follow the same split as steady-state flow.
T+6 to T+10 seconds. If a standby pump is configured to auto-start, it typically reaches rated speed in this window, but the network does not return to its prior flow distribution instantly. Pressure and flow oscillate as the system settles, and cold plates on the branches that lost the most flow are still running hotter than the network average when the standby pump reaches full speed.
Why this matters for design
Junction temperature during those ten seconds is a function of local flow, not average network flow. A design that verifies "N+1 pump redundancy restores full flow within 10 seconds" without verifying the transient distribution during those 10 seconds is verifying the wrong thing.
Modeling this sequence requires a transient solver on the full network topology, not a static pressure-drop calculation with a redundancy factor bolted on. That is what a pump-trip scenario should mean in a design review: not "does a second pump exist," but "what does every cold plate see, second by second, until the network is back at steady state."