Contact us

Cooling custom silicon at scale

Posted on
October 8, 2026

Cooling starts at the silicon

The cooling requirements of an AI accelerator are not defined by its power consumption. A 1,500-watt package is not simply a 1,500-watt heat load. Engineers must understand where those watts originate and how heat travels from the silicon through the package because two processors consuming equal amounts of power may have very different cooling requirements. Silicon thermal signatures depend on die configuration, the location and intensity of compute cores, hot spots, high-bandwidth memory placement, package construction, thermal resistance, and allowable operating temperatures. Modern AI accelerators may also combine multiple compute dies, memory stacks, interposers, and other components within a single package, and different components may have different thermal limits.

McKinsey notes that GPU and custom silicon architectures establish power density, thermal, and interconnect requirements years in advance, making cooling part of the compute architecture. Cooling development can begin before production silicon is available. Thermal test vehicles (TTVs) modeled on the expected characteristics of a future processor can simulate its heat distribution and other relevant package characteristics, allowing engineers to physically test and refine cooling solutions while silicon development is still underway. That puts cooling and silicon development on parallel paths.

Design the cold plate around the processor

Cooling custom silicon at scale

For many advanced CPUs, GPUs, ASICs, networking devices, and AI accelerators, the cold plate is the first place where silicon-specific cooling becomes real hardware. It has to match the thermal profile of the package, but it also has to work within the mechanical, fluid, and integration constraints of the platform.

That means the design is not just about removing more heat. It is about removing heat efficiently, keeping thermal resistance low, managing pressure drop, and fitting into a system that can be built, serviced, and scaled.

Once the package layout and high-heat-flux areas are understood, the internal cold plate architecture can be tuned to the processor instead of treating the package as a uniform heat source. Flow paths, microconvection impingement features, materials, mounting approach, and thermal interfaces are used to create a solution that addresses the areas that need the most cooling, or specific types of cooling. That is especially important for custom silicon, where die placement, memory location, package construction, and allowable temperatures can vary significantly from one design to the next. JetCool and Broadcom, for example, recently announced a direct-to-chip cooling solution for next-generation custom AI XPUs designed to support sustained multi-kilowatt ASIC operation at heat flux levels reaching 4 W/mm2.

That kind of expert design work is where customization delivers value. The cold plate is engineered around the processor’s thermal and mechanical requirements, while also accounting for the operating conditions of the rack, CDU, and facility water system. This holistic silicon-to-facility approach to cooling design helps data centers improve power and water efficiency, increase density, and maintain performance under peak workloads.

Standardize the boundaries. Differentiate within them.

At hyperscale, small improvements in performance per watt, utilization, or workload economics can multiply across large fleets. That helps explain the investment in specialized silicon. McKinsey estimates that customized chips could reduce cost per token by 70 percent to 80 percent for the inference workloads they are designed to handle.

The challenge is that processor-specific optimization cannot turn the rest of the infrastructure into a one-off design. At the silicon, cooling can be tuned to the heat map, package design, and thermal limits of a specific processor. Farther from the chip, common approaches to fluid connections, operating conditions, manifolds, racks, CDUs, facility water systems, and telemetry allow different types and generations of compute to coexist in production, increasing density and simplifying deployment.

Standardization does not eliminate customization. It defines where customization belongs. With expert design, the most processor-specific work can stay close to the silicon while the rack and facility infrastructure remain consistent to support repeatable deployment at scale. As custom compute becomes more common, that boundary becomes more important.

Customize where physics and performance demand it.

Standardize where repeatability and scale depend on it.

Turn custom cooling into repeatable infrastructure

There is a big difference between engineering a cold plate that works and manufacturing a thermal system that can be reproduced at full scale. Internal channel geometry, flatness, surface finish, joining processes, material compatibility, cleanliness, pressure integrity, thermal-interface consistency, and flow distribution all affect final cold plate performance. In development, small variations can often be studied, adjusted, or managed by engineering. At production scale, those same variations have to be controlled across thousands of servers and potentially hundreds of thousands of processors.

That makes manufacturing capability part of the cooling solution, not a handoff after the design is finished. Leak and pressure testing, flow characterization, thermal validation, cleanliness controls, and process monitoring all need to scale with production. They have to provide confidence in every unit without slowing the line or turning validation into a bottleneck.

For processor-specific cooling, that gap between prototype and production is especially important. A cold plate tuned to a particular heat map has to reproduce the intended flow and cooling behavior consistently as volume increases, while still connecting cleanly to manifolds, racks, CDUs, and facility water systems. The custom silicon provider needs a partner that can do both: engineer the cooling architecture around the processor and build it with the process control, repeatability, and scale required for deployment.

Engineer the system from chip to facility

Custom silicon cooling designed at the chip level

AI infrastructure now depends on decisions made across the full stack: silicon, packaging, power delivery, cooling, mechanical design, manufacturing, and facilities. That is why custom cooling cannot be designed around the processor alone. It has to support the thermal needs of the chip, connect cleanly into the tray, rack, CDU, and facility water system, and be ready for the next generation of compute without forcing a redesign of the surrounding infrastructure. TTVs and early physical testing help cooling development move in parallel with silicon design, while manufacturing engineering ensures the solution can be built repeatedly at production scale.

Our AI infrastructure capabilities span processor-specific cold plates and direct-to-chip liquid cooling through manifolds, CDUs, IT racks, embedded and critical power, and integrated data center infrastructure. Combined with systems engineering and advanced manufacturing at scale, that breadth creates a path from a cooling solution tuned to the heat map of a particular processor to repeatable rack- and facility-level infrastructure and high-volume deployment.

Learn more about our cooling solutions for custom silicon at the 2026 OCP Global Summit