CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
AuthorsKaustav Goswami, Maryam Babaie, Hoa Nguyen, Venkatesh Akella, Jason Lowe-Power
Resources
This paper presents CXL-ClusterSim, a new simulation framework for studying how pooled, disaggregated memory over CXL could make large-scale AI and cloud systems more efficient.
Key results
The simulated remote link bandwidth reported by STREAM is within 0.99% of the baseline measurements.
SST’s memory controller statistics are reported to be within 0.1% of the baseline.
In the 8-host STREAM study, interleaving data across local and remote memory reduced effective bandwidth to 6.45 GiB/s per node.
The same STREAM study reports 11.4 GiB/s for strictly local memory.
Adding 170 ns CXL latency reduced STREAM bandwidth by 8.95%.
Adding 250 ns CXL latency reduced STREAM bandwidth by 29%.
What the paper found
Kaustav Goswami, Maryam Babaie, Hoa Nguyen, Venkatesh Akella, and Jason Lowe-Power at the University of California, Davis present CXL-ClusterSim, a full-system simulation framework for CXL-based disaggregated memory that combines gem5’s high-fidelity, OS-visible execution with SST’s MPI-parallel scaling. The novelty is a two-phase workflow: gem5 fast-forwards each host through Linux boot and application initialization, checkpoints the region of interest, then SST restores synchronized timing-accurate execution across multiple hosts and a remote memory blade. The framework models both CXL-style memory pooling, exposed as NUMA nodes, and memory sharing, exposed as DAX /dev/dax mappings for direct byte-addressable access. Validation shows the simulated remote link reports STREAM bandwidth within 0.99% of baseline measurements, and SST’s memory controller statistics are within 0.1%. In an 8-host STREAM study, pinning data to the remote pool reduced effective bandwidth to 6.45 GiB/s per node under interleaving, versus 11.4 GiB/s for local memory, while adding 170 ns and 250 ns CXL latency reduced bandwidth by 8.95% and 29%, respectively. For pooling, NAS Parallel Benchmarks class D showed that larger remote fractions sharply reduce IPC, with mg falling to 0.38 relative IPC when 52% of its 27 GiB working set spilled remote. For sharing, GAPBS kernels on a single shared graph achieved 31.8% of memory instructions from the remote blade, but pointer-chasing kernels suffered the largest slowdowns. The authors position CXL-ClusterSim as an open-source platform for exploring hardware/software co-design beyond NUMA emulation and note future work on hot-plugging, coherence, and switch modeling.
Original abstract
Large-scale AI training and inference require hundreds of gigabytes to terabytes of DRAM with high peak to average utilization ratios, resulting in overprovisioning. In cloud computing, DRAM constitutes a significant share of the cost. Yet, as shown by recent articles, DRAM is heavily under utilized. Memory disaggregation is a solution to both these problems. With the advent of the CXL protocol, there is renewed interest in designing and optimizing computing systems with disaggregated memory. However, at present, there are limited simulation tools available for exploring the design space and evaluating the performance tradeoffs in computer systems with disaggregated memory. In this paper, we propose CXL-ClusterSim, a full-system modeling and simulation framework by combining the gem5 simulator for fidelity, with the Structural Simulation Toolkit (SST) for parallel simulation. We outline the challenges in creating this simulation infrastructure and present a design that is scalable, flexible, and reasonably fast to help computer architects to explore the design space of CXL-based disaggregated memory and identify new opportunities for hardware/software codesign and performance optimization.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.