Hardware Mechanisms to Dynamically Throttle AI Performance
AuthorsHaiyue Ma, Lauren Malek, Joseph Forzani, David Wentzlaff
Resources
The paper proposes GPU hardware controls that can dynamically turn down an AI system’s performance as a last-resort safety mechanism.
Key results
Performance reduction at one-eighth resource availability.
Largest reported prefill sensitivity among evaluated clock knobs.
Flip-flop budget; each mechanism uses less than this amount.
Approximate cycles to stabilize after the deepest cut.
Approximate cycles to stabilize at the one-eighth cut.
What the paper found
Princeton researchers Haiyue Ma, Lauren Malek, Joseph Forzani, and David Wentzlaff propose hardware-enforced, runtime control for AI capability that complements software safeguards used by systems from OpenAI and Anthropic. Instead of a disruptive full-chip shutdown, their GPU memory-system throttlers selectively reduce L2 cache associativity, L2 latency, L2 response bandwidth, or shared-memory port access, using cache-way masking, latency insertion, credit-based rate limiting, and bank arbitration. In AccelSim simulations of an NVIDIA A100 using NVIDIA CUTLASS kernels representative of DeepSeek-V3, Llama-3-70B, and Mixtral-8x7B, L2 latency reduced decode performance by 80 percent at one-eighth resource availability, while interconnect-clock throttling reduced prefill performance by 87 percent but affected workloads more globally. The proposed mechanisms require less than 10K flip-flops each, switch throttling levels in one cycle, and stabilize in roughly 5K to 80K cycles; shared-memory throttling settled in roughly 5K cycles, whereas L2 associativity required roughly 80K cycles because masked cache lines are evicted gradually. Multi-knob combinations can amplify slowdowns, particularly pairing L2-size reduction with DRAM throttling during prefill. Tests with six CUDA Samples workloads indicate that memory-side controls are substantially more selective than disabling compute cores, preserving capacity for non-AI workloads while imposing a physical limit that software, kernel rewriting, or model behavior cannot bypass.
Original abstract
As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. Existing software safeguards impose only behavioral constraints that can potentially be bypassed by sufficiently intelligent models. While hardware-level safety enforcement has been recognized as an essential last line of defense, few mechanisms have been proposed beyond policy regulations on unauthorized accesses or coarse full-chip shutdown. What is missing is a fine-grained, dynamic intervention mechanism at the architecture level. In this paper, we introduce a set of microarchitecture knobs which dynamically control the available hardware resources to limit AI performance at runtime. We evaluate candidate knobs spanning the GPU memory subsystem, across capacity, bandwidth, latency and frequency dimensions, and narrow down to four strong candidates: L2 size, L2 latency, L2 bandwidth, and shared memory port access rate. To minimize new logic and extra design cost, we build all four mechanisms from well-established microarchitectural primitives: cache way masking, credit-based rate limiting, latency insertion, and bank arbitration. We show that these knobs achieve high performance sensitivity (up to 80% performance cut at 1/8 resource availability), negligible implementation cost (<~10K flip flops), fast stabilization after dynamic throttling (5-80K cycles), and minimal collateral impact on the rest of the chip. Further, multi-knob analysis reveals combinations of knobs that amplify the performance degradation beyond the effect of each knob individually, which enables a broader range of performance targets.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.