NTH

Hardware Mechanisms to Dynamically Throttle AI Performance

AuthorsHaiyue Ma, Lauren Malek, Joseph Forzani, David Wentzlaff

July 24, 2026 2 min read
Watch on YouTube
The one-line take

The paper proposes GPU hardware controls that can dynamically turn down an AI system’s performance as a last-resort safety mechanism.

Key results

80%
Decode L2 latency performance cut

Performance reduction at one-eighth resource availability.

87%
Prefill interconnect-clock performance cut

Largest reported prefill sensitivity among evaluated clock knobs.

10K
Per-throttler hardware overhead

Flip-flop budget; each mechanism uses less than this amount.

5K
Shared-memory throttler stabilization

Approximate cycles to stabilize after the deepest cut.

80K
L2-associativity stabilization

Approximate cycles to stabilize at the one-eighth cut.

What the paper found

Princeton researchers Haiyue Ma, Lauren Malek, Joseph Forzani, and David Wentzlaff propose hardware-enforced, runtime control for AI capability that complements software safeguards used by systems from OpenAI and Anthropic. Instead of a disruptive full-chip shutdown, their GPU memory-system throttlers selectively reduce L2 cache associativity, L2 latency, L2 response bandwidth, or shared-memory port access, using cache-way masking, latency insertion, credit-based rate limiting, and bank arbitration. In AccelSim simulations of an NVIDIA A100 using NVIDIA CUTLASS kernels representative of DeepSeek-V3, Llama-3-70B, and Mixtral-8x7B, L2 latency reduced decode performance by 80 percent at one-eighth resource availability, while interconnect-clock throttling reduced prefill performance by 87 percent but affected workloads more globally. The proposed mechanisms require less than 10K flip-flops each, switch throttling levels in one cycle, and stabilize in roughly 5K to 80K cycles; shared-memory throttling settled in roughly 5K cycles, whereas L2 associativity required roughly 80K cycles because masked cache lines are evicted gradually. Multi-knob combinations can amplify slowdowns, particularly pairing L2-size reduction with DRAM throttling during prefill. Tests with six CUDA Samples workloads indicate that memory-side controls are substantially more selective than disabling compute cores, preserving capacity for non-AI workloads while imposing a physical limit that software, kernel rewriting, or model behavior cannot bypass.

Original abstract

As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. Existing software safeguards impose only behavioral constraints that can potentially be bypassed by sufficiently intelligent models. While hardware-level safety enforcement has been recognized as an essential last line of defense, few mechanisms have been proposed beyond policy regulations on unauthorized accesses or coarse full-chip shutdown. What is missing is a fine-grained, dynamic intervention mechanism at the architecture level. In this paper, we introduce a set of microarchitecture knobs which dynamically control the available hardware resources to limit AI performance at runtime. We evaluate candidate knobs spanning the GPU memory subsystem, across capacity, bandwidth, latency and frequency dimensions, and narrow down to four strong candidates: L2 size, L2 latency, L2 bandwidth, and shared memory port access rate. To minimize new logic and extra design cost, we build all four mechanisms from well-established microarchitectural primitives: cache way masking, credit-based rate limiting, latency insertion, and bank arbitration. We show that these knobs achieve high performance sensitivity (up to 80% performance cut at 1/8 resource availability), negligible implementation cost (<~10K flip flops), fast stabilization after dynamic throttling (5-80K cycles), and minimal collateral impact on the rest of the chip. Further, multi-knob analysis reveals combinations of knobs that amplify the performance degradation beyond the effect of each knob individually, which enables a broader range of performance targets.

Read the original paper

More in AI Hardware

Browse all 34 papers →