NTH

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU

AuthorsReese Levine, Rithik Sharma, Nikhil Jain, Abhijit Ramesh, Zheyuan Chen, Neha Abbas, James Contini, Tyler Sorensen

May 25, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows how to run large language models efficiently in the browser with WebGPU, cutting memory use and boosting speed across many devices and quantization formats.

Key results

29–33%
Memory usage reduction

LlamaWeb reduces peak browser memory versus WebLLM and Transformers.js across 16 devices from 8 vendors and 10 models.

45–69%
Decode throughput improvement

Across four GPUs, LlamaWeb increases decode throughput over WebLLM and Transformers.js.

up to 59%
Apple M4 Pro prefill speedup vs Vulkan

On Apple M4 Pro, the WebGPU backend is reported to be faster than Vulkan during prefill.

23%
Intel Arc B580 prefill speedup vs SYCL

With WebGPU checks disabled, LlamaWeb outperforms native SYCL on Intel Arc B580 during prefill.

23
Supported weight formats

The backend supports 23 data formats / weight formats for inference in WebGPU, extending llama.cpp quantization support.

What the paper found

Llamas on the Web introduces LlamaWeb, a WebGPU backend for llama.cpp that targets browser-based LLM inference with static memory planning, performance-portable kernels, and broad quantization support. The key novelty is a design that eliminates dynamic GPU allocation by preallocating all buffers, including a fixed-slot parameter arena and intermediate memory for FlashAttention and FlashDecoding, while streaming GGUF weights asynchronously from the browser’s Origin Private File System without materializing them in the WebAssembly heap. Across 16 devices from 8 vendors and 10 models, LlamaWeb reduces peak browser memory by 29–33 percent relative to WebLLM and Transformers.js, and avoids the memory growth and tab crashes observed in those systems. For performance, it uses a templated shader library that specializes kernels by device features such as subgroups and subgroup matrix operations, and by workload parameters such as tile size, with tuning derived from an exhaustive sweep of thousands of configurations. On four GPUs, this portability layer improves decode throughput by 45–69 percent over WebLLM and Transformers.js, and on some systems it matches or beats native backends: on Apple M4 Pro it is up to 59 percent faster than Vulkan during prefill, and on Intel Arc B580 it outperforms native SYCL by 23 percent in prefill with WebGPU checks disabled. The paper also extends llama.cpp’s quantization ecosystem by implementing templated dequantization for 23 weight formats, including legacy q4_0 and q8_0, K-quants such as q4_k_m, I-quants, and the new q1_0 format, enabling one kernel family to support many compressed models without format-specific rewrites.

Original abstract

Running language models in the browser presents a unique opportunity to build efficient, private, and portable AI applications, but requires contending with constrained memory availability and heterogeneous hardware targets. To realize this opportunity, we present Llamas on the Web (LlamaWeb), a WebGPU backend for llama.cpp that enables memory-efficient and performance-portable LLM inference across a wide range of model weight formats in the browser. Our design significantly reduces memory overhead through static memory planning and efficient model loading, addresses cross-device variability through a tunable kernel library, and introduces templated GPU kernels that support performant implementations of numerous quantization formats, enabling broad model support and extensibility to new formats. We evaluate LlamaWeb on 16 devices from 8 vendors, collecting data from 10 language models and four model weight formats. We compare LlamaWeb against existing browser-based LLM frameworks and find that LlamaWeb requires 29-33% less memory across several combinations of device, browser, and operating system. We also evaluate LlamaWeb's performance against these frameworks and find that it increases decode throughput by 45-69% across four GPUs from separate vendors. In addition, we compare LlamaWeb's performance against other llama.cpp backends, where it is competitive with and even beats vendor-specific backend performance on some devices.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis