NTH

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

AuthorsZeyu Cao, Xuan Guo, Cheng Zhang, Cheuk Hang Lau, Ilia Shumailov, Yiren Zhao

August 22, 2026 2 min read
Watch on YouTube
The one-line take

A 128-GPU cluster assembled from discarded hardware can serve LLaMA-70B cheaply, but its environmental value depends heavily on electricity costs and carbon intensity.

Key results

22K
DumpsterCluster capital cost

Total cost of the 128-GPU second-hand system, versus $600K for an 8-GPU NVIDIA B200 system.

80
Pipeline stages

Maximum device-level pipeline depth supported for LLaMA-70B inference.

223K
LLaMA 3.1-8B throughput

Tokens per second achieved by 128 V100 GPUs on a prefill-heavy workload.

1.5K
LLaMA 3.1-70B throughput

Tokens per second achieved by 128 V100 GPUs on a prefill-heavy workload.

3.7
8B carbon increase

Approximate carbon-emissions multiplier per token versus newer hardware under grid-average electricity.

13.8%
Hardware-replacement incident share

Share of 58 incidents requiring physical hardware replacement during one year.

What the paper found

DumpsterCluster asks whether retired hardware can serve modern AI economically, building a 128-GPU cluster entirely from second-hand components and operating it for one year. Using NVIDIA V100 cards that now cost about $60 each, the complete system costs $22K, compared with $600K for an 8-GPU NVIDIA B200 system. Its custom Rust serving engine replaces the tensor-parallel approach common in vLLM with device-level, pipeline-first parallelism, asynchronously moving activations through CPU-managed communication and supporting up to 80 pipeline stages. This lets the cluster serve Meta’s LLaMA 3.1-70B, which cannot fit on a single V100, while delivering 223K TPS on LLaMA 3.1-8B in a prefill-heavy workload and 1.5K TPS on LLaMA 3.1-70B. The economic advantage comes with an environmental trade-off: under grid-average electricity, reused V100 systems produce about 3.7 times more carbon per token for 8B inference and 41 times more for 70B inference than newer accelerators. Reliability was manageable after stress-testing and redundancy; only 13.8% of 58 service incidents required hardware replacement. The study concludes that second-hand GPUs are viable primarily for inference, especially where electricity is inexpensive and low-carbon, but they are poorly suited to bandwidth-intensive training.

Original abstract

As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM inference, and under what conditions such repurposing is economically viable and environmentally sustainable. We physically built a 128-GPU DumpsterCluster from scratch using only second-hand components and ran it for one year. At current market prices (\$22K for the DumpsterCluster vs. \$600K for an 8-GPU B200 system), the economic advantages are substantial. Through pipeline-parallel optimizations, our V100 based DumpsterCluster achieves competitive LLaMA-70B throughput, validating production viability. However, our deployment reveals critical context dependencies. Older GPUs consume significantly more energy per token, making total cost of ownership favorable only in regions with inexpensive electricity. Under grid-average carbon intensity, second-hand systems can produce approximately 4x higher total carbon emissions per token for 8B models, and over 40x for 70B models, compared to current-generation hardware. These findings show that GPU afterlife is not universally sustainable - hardware repurposing must be strategically coupled with low carbon energy sources. When deployed in regions with favourable energy economics and clean electricity, second-hand GPUs offer a viable pathway for expanding AI capacity while advancing affordability, energy security, and environmental responsibility.

Read the original paper

More in AI Hardware

Browse all 34 papers →