NTH

CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST

AuthorsKaustav Goswami, Maryam Babaie, Hoa Nguyen, Venkatesh Akella, Jason Lowe-Power

June 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper presents CXL-ClusterSim, a new simulation framework for studying how pooled, disaggregated memory over CXL could make large-scale AI and cloud systems more efficient.

Key results

0.99%
STREAM validation error

The simulated remote link bandwidth reported by STREAM is within 0.99% of the baseline measurements.

0.1%
SST controller error

SST’s memory controller statistics are reported to be within 0.1% of the baseline.

6.45 GiB/s
Remote bandwidth per node

In the 8-host STREAM study, interleaving data across local and remote memory reduced effective bandwidth to 6.45 GiB/s per node.

11.4 GiB/s
Local bandwidth baseline

The same STREAM study reports 11.4 GiB/s for strictly local memory.

8.95%
Bandwidth drop with 170 ns latency

Adding 170 ns CXL latency reduced STREAM bandwidth by 8.95%.

29%
Bandwidth drop with 250 ns latency

Adding 250 ns CXL latency reduced STREAM bandwidth by 29%.

What the paper found

Kaustav Goswami, Maryam Babaie, Hoa Nguyen, Venkatesh Akella, and Jason Lowe-Power at the University of California, Davis present CXL-ClusterSim, a full-system simulation framework for CXL-based disaggregated memory that combines gem5’s high-fidelity, OS-visible execution with SST’s MPI-parallel scaling. The novelty is a two-phase workflow: gem5 fast-forwards each host through Linux boot and application initialization, checkpoints the region of interest, then SST restores synchronized timing-accurate execution across multiple hosts and a remote memory blade. The framework models both CXL-style memory pooling, exposed as NUMA nodes, and memory sharing, exposed as DAX /dev/dax mappings for direct byte-addressable access. Validation shows the simulated remote link reports STREAM bandwidth within 0.99% of baseline measurements, and SST’s memory controller statistics are within 0.1%. In an 8-host STREAM study, pinning data to the remote pool reduced effective bandwidth to 6.45 GiB/s per node under interleaving, versus 11.4 GiB/s for local memory, while adding 170 ns and 250 ns CXL latency reduced bandwidth by 8.95% and 29%, respectively. For pooling, NAS Parallel Benchmarks class D showed that larger remote fractions sharply reduce IPC, with mg falling to 0.38 relative IPC when 52% of its 27 GiB working set spilled remote. For sharing, GAPBS kernels on a single shared graph achieved 31.8% of memory instructions from the remote blade, but pointer-chasing kernels suffered the largest slowdowns. The authors position CXL-ClusterSim as an open-source platform for exploring hardware/software co-design beyond NUMA emulation and note future work on hot-plugging, coherence, and switch modeling.

Original abstract

Large-scale AI training and inference require hundreds of gigabytes to terabytes of DRAM with high peak to average utilization ratios, resulting in overprovisioning. In cloud computing, DRAM constitutes a significant share of the cost. Yet, as shown by recent articles, DRAM is heavily under utilized. Memory disaggregation is a solution to both these problems. With the advent of the CXL protocol, there is renewed interest in designing and optimizing computing systems with disaggregated memory. However, at present, there are limited simulation tools available for exploring the design space and evaluating the performance tradeoffs in computer systems with disaggregated memory. In this paper, we propose CXL-ClusterSim, a full-system modeling and simulation framework by combining the gem5 simulator for fidelity, with the Structural Simulation Toolkit (SST) for parallel simulation. We outline the challenges in creating this simulation infrastructure and present a design that is scalable, flexible, and reasonably fast to help computer architects to explore the design space of CXL-based disaggregated memory and identify new opportunities for hardware/software codesign and performance optimization.

Read the original paper

More in AI Hardware

Browse all 34 papers →