MaxKernel: Agentic Kernel Generation for TPUs
AuthorsShangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran
MaxKernel uses collaborating AI agents and compiler feedback to automatically discover high-performance TPU kernels that can rival expert optimization.
Key results
Diverse TPU kernel tasks used for evaluation
Geometric-mean speedup over the XLA baseline
MaxKernel parallel search across eight kernels
Acceleration over the JAX implementation
Sparse Attention acceleration over JAX
What the paper found
MaxKernel is a multi-agent system for generating and optimizing TPU kernels in JAX and Pallas, using compiler feedback and XProf hardware traces rather than relying on static, zero-shot code generation. Its specialized agents plan algorithms, write and repair kernels, synthesize tests, autotune tile configurations, verify numerical equivalence, and profile bottlenecks. The framework supports three workflows: Human-in-the-Loop collaboration, a closed-loop Autonomous agent, and Graph-Based Autonomous Search using parallel or beam exploration to avoid local optima. Powered in these experiments by Gemini 3.1 Pro and evaluated on JAXBench’s 50 diverse TPU kernel tasks, MaxKernel’s parallel search achieved a 1.58× geometric-mean speedup over the XLA baseline, with 50/50 compilation and correctness. On eight production kernels with human-written references, it reached a 2.32× geometric-mean speedup versus 2.02× for hand-tuned implementations. The gains extend to real workloads from models including Qwen3-Next and DeepSeek-V4: generated kernels accelerated the Qwen3-Next Gated DeltaNet training step by 4.70× and DeepSeek-V4 Sparse Attention prefill by 7.85× over JAX. MaxKernel also reduced Multi-Head Latent Attention latency by 8.68 percent and resolved Ragged Paged Attention crashes by inserting protective ALU clamps, showing that agentic search can address both performance and reliability in low-level accelerator programming.
Original abstract
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.