BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
AuthorsZili Zhang, Chengxu Yang, Shenglong Zhang, Chenyu Wang, Yufan Zhang, Tuo Dai, Zhouyang Li, Yuhong Ge, Chao Jin, Xin Jin, Yuliang Liu
Resources
BigMac is a new multimodal LLM training pipeline that claims to break the usual tradeoff between speed and memory, making training both faster and more memory-friendly.
Key results
On MLLM-Understanding, BigMac improves training iteration time over Optimus.
On MLLM-Understanding, BigMac improves training iteration time over Megatron-DistTrain.
On MLLM-Generation, BigMac improves training iteration time over Megatron-DistTrain.
The production deployment result is from an internal multimodal LLM training run at this parameter scale.
The production deployment ran on this many Hopper GPUs.
Across the first 128 iterations in the bitwise alignment study, the average absolute loss difference stayed below this threshold.
What the paper found
BigMac introduces a dependency-safe nested pipeline for multimodal LLM training that breaks the usual tradeoff between throughput and activation memory. Existing compute-efficient systems such as Optimus decouple the encoder and generator from the LLM pipeline, but this forces retention of encoder activations across the full pipeline and can also задержle LLM activations until generation completes; memory-efficient systems such as Megatron-DistTrain integrate all modules into one pipeline, but then model and data heterogeneity create bubbles and stall the slowest stage. BigMac keeps the user-specified LLM schedule, including interleaved 1F1B, and nests encoder and generator microbatch units only at dependency-safe points, so the encoder and generator each require only O(1) activation memory while the LLM backbone remains unchanged. It also replaces rigid FSDP all-gather with a one-sided NVSHMEM pull to remove synchronization barriers, and supports distinct context-parallel groups for encoder and LLM. On a 128-GPU Hopper cluster, BigMac speeds up MLLM-Understanding with Qwen3-30B-A3B plus a 1.3B ViT by 1.08×–1.9× over the baselines, and MLLM-Generation with the same backbone plus a 20B MMDiT by 1.5×–1.9×, while keeping peak memory flat as batch size increases and avoiding the out-of-memory failures seen in Optimus. In production on a 345B-parameter MLLM across 1,536 Hopper GPUs and more than 18K iterations, it preserved stable loss dynamics, with average loss deviation below 10^-3 versus Megatron-LM in controlled tests.
Original abstract
Training multimodal large language models (MLLMs) is challenged by both model and data heterogeneity. Existing systems redesign the training pipeline to address these challenges, but remain bound by a Pareto frontier between compute and memory efficiency, improving one only at the expense of the other. We present BigMac, a new training pipeline for multimodal LLMs. The core idea of BigMac is to elegantly nest the encoder and generator computation into the original LLM pipeline, forming a dependency-safe nested pipeline structure. With this design, BigMac reduces the activation memory complexity of the encoder and generator to O(1) while keeping the activation memory complexity of the LLM unchanged. At the same time, it achieves the same computational efficiency as the idealized setting with unlimited memory. As a result, BigMac breaks the Pareto frontier between computational efficiency and memory usage, enabling simultaneous optimization of both computation and memory in MLLM training. We evaluate BigMac on multiple MLLMs and training workloads. Experimental results show that BigMac achieves a 1.08$\times$-1.9$\times$ training speedup over baseline systems while maintaining stable memory usage as batch size increases.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.