Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
AuthorsArchitect Labs
Resources
Redwood is an AI-designed accelerator that claims to go from specification to verified silicon-ready hardware and efficient model execution in just two weeks.
Key results
Complete accelerator design, verification, firmware, and kernel generation time in weeks
Code and functional coverage reached by every block
Maximum hours to regenerate, reverify, and redeploy architectural changes
Average Qwen3-0.6B decode throughput in tokens/s on the projected ASIC
Total chip-side power in watts for the projected Samsung 8 nm-class implementation
Projected improvement over the NVIDIA Jetson Orin Nano baseline
What the paper found
Redwood demonstrates an end-to-end AI-generated accelerator designed, verified, programmed, and deployed from a high-level specification in under 2 weeks. Its tile-based spatial-dataflow architecture combines RISC-V control cores, systolic GEMM and GEMV engines, SIMD units, banked scratchpads, and DMA-driven memory movement, with hardware-scheduled kernels for transformer operations such as GEMM and FlashAttention; an emulated-softmax technique from FlashAttention-4 reuses SIMD hardware to avoid a dedicated unit. Redwood Nano uses a 2×2 tile array and runs Qwen3-0.6B on an AMD Versal FPGA, while the projected Samsung 8 nm-class implementation targets multi-billion-parameter models including Llama and Qwen. The automated flow generated RTL, UVM verification environments, formal proofs, firmware, and kernels, reaching 95% code and functional coverage for every block. Architectural revisions were regenerated, reverified, and redeployed in under 48 hours. Against an NVIDIA Jetson Orin Nano running the same Qwen model, the projected ASIC reaches 49 tokens/s at 1.335 W, compared with 28 tokens/s at 2.59 W, yielding a reported 3.4x performance-per-watt improvement. The measured FPGA system achieved 12.1 average tokens/s, and Qwen3 running on Redwood helped identify optimizations for a subsequent accelerator generation, providing an early example of hardware-enabled recursive self-improvement.
Original abstract
Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.