Chiplet3D: Pin- and Thermal-Aware 3D Chiplet Floorplanning via Convolution-Embedded MILP
AuthorsShuo Ren, Libo Shen, Yaohui Han, Rongliang Fu, Junying Huang, Bei Yu, Tsung-Yi Ho
Resources
Chiplet3D uses a thermally informed MILP to arrange 3D chiplets and pins for shorter connections while substantially reducing hotspots.
Key results
Average reduction versus state-of-the-art baselines on ICCAD’24 ATPlace.
Maximum observed wirelength reduction.
Maximum reduction in peak temperature, measured in degrees C.
Maximum reduction in thermal non-uniformity.
Cases won out of 10 for bottom-die peak temperature.
Cases won out of 10 for thermal uniformity.
What the paper found
Chiplet3D, from Shuo Ren and colleagues, is a two-die 3D chiplet floorplanner that addresses a central limitation of stacked designs: heat trapped between active dies, as seen in products such as AMD 3D V-Cache, TSMC SoIC, and Intel Foveros. Its first novelty is pin-aware half-perimeter wirelength: instead of connecting block centers, the MILP precomputes exact pin coordinates under four rotations and two flips, allowing orientation choices that align real terminals. Its second novelty is a coarse convolutional thermal field embedded directly in the MILP. Calibrated point-spread kernels approximate lateral and vertical heat conduction on a grid, while a two-stage flow first creates a thermal skeleton and then uses dynamic thermal elasticity to recover shorter wiring without violating thermal constraints. Evaluated on the ICCAD’24 ATPlace benchmarks, with final temperatures verified by golden 3D-ICE simulations, Chiplet3D reduces wirelength by 39% to 43% on average and by up to 62% in the best case. It lowers peak temperature by up to 45.9 degrees C and thermal non-uniformity by up to 56% versus state-of-the-art baselines. Across 10 cases, it achieves the lowest bottom-die peak temperature in 8 cases and the best thermal uniformity in 9, while producing the shortest average pin-aware wirelength. The results show that exact pin alignment and solver-integrated thermal spreading create a stronger wirelength–temperature Pareto frontier than center-based or post-placement thermal estimation.
Original abstract
As traditional Moore's Law scaling slows down, 3D-ICs stack multiple active dies vertically to sustain performance scaling. However, this vertical stacking traps heat inside, making temperature a design concern. Although we can fix thermal issues at different design steps, floorplanning is the earliest and most cost-effective stage to solve it. Previous methods handle this by assuming wires connect to block centers and estimating temperature through simplistic power-based calculations, but these assumptions mislead their wirelength optimization and leave hotspots unresolved. To address these limitations, we present Chiplet3D, a pin- and thermal-aware floorplanner for two-die 3D-ICs. To achieve pin-awareness, it supports all four rotations and two flips, measuring wirelength from exact pin locations so the solver can flip or rotate blocks to pull connected pins closer. On the thermal side, Chiplet3D replaces the inaccurate power-based metrics of prior work with a fast, coarse convolution field embedded directly in a mixed-integer linear program (MILP) to accurately track the true 3D heat spread. We evaluate Chiplet3D on the ICCAD'24 ATPlace benchmarks, validating every temperature with a golden 3D-ICE simulation. Chiplet3D reduces wirelength by 39\%--43\% on average (and up to 62\% in the best case), while lowering peak temperatures by up to 45.9$^\circ$C and reducing thermal non-uniformity by up to 56\% compared to the SOTA baselines. Overall, these results demonstrate that by co-optimizing pin alignment and thermal fields, Chiplet3D establishes a stronger Pareto frontier between thermal-aware layout and interconnect efficiency.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.