NTH

MaxKernel: Agentic Kernel Generation for TPUs

AuthorsShangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran

September 8, 2026 2 min read
Watch on YouTube
The one-line take

MaxKernel uses collaborating AI agents and compiler feedback to automatically discover high-performance TPU kernels that can rival expert optimization.

Key results

50
JAXBench task count

Diverse TPU kernel tasks used for evaluation

1.58
JAXBench parallel-search speedup

Geometric-mean speedup over the XLA baseline

2.32
Production-kernel geometric speedup

MaxKernel parallel search across eight kernels

4.70
Qwen3-Next training-step speedup

Acceleration over the JAX implementation

7.85
DeepSeek-V4 prefill speedup

Sparse Attention acceleration over JAX

What the paper found

MaxKernel is a multi-agent system for generating and optimizing TPU kernels in JAX and Pallas, using compiler feedback and XProf hardware traces rather than relying on static, zero-shot code generation. Its specialized agents plan algorithms, write and repair kernels, synthesize tests, autotune tile configurations, verify numerical equivalence, and profile bottlenecks. The framework supports three workflows: Human-in-the-Loop collaboration, a closed-loop Autonomous agent, and Graph-Based Autonomous Search using parallel or beam exploration to avoid local optima. Powered in these experiments by Gemini 3.1 Pro and evaluated on JAXBench’s 50 diverse TPU kernel tasks, MaxKernel’s parallel search achieved a 1.58× geometric-mean speedup over the XLA baseline, with 50/50 compilation and correctness. On eight production kernels with human-written references, it reached a 2.32× geometric-mean speedup versus 2.02× for hand-tuned implementations. The gains extend to real workloads from models including Qwen3-Next and DeepSeek-V4: generated kernels accelerated the Qwen3-Next Gated DeltaNet training step by 4.70× and DeepSeek-V4 Sparse Attention prefill by 7.85× over JAX. MaxKernel also reduced Multi-Head Latent Attention latency by 8.68 percent and resolved Ragged Paged Attention crashes by inserting protective ALU clamps, showing that agentic search can address both performance and reliability in low-level accelerator programming.

Original abstract

Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.

Read the original paper

More in AI Hardware

Browse all 34 papers →