NTH

Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?

AuthorsNishant Aggarwal, Ayushi Dubal, Sreeraj Kannakarankodi, Ian McDougall, Adarsh Mittal, Vishnu Ramadas, Noah Scott, Ranganath Selagamsetty, Weichu Yang, Karthikeyan Sankaralingam

July 17, 2026 2 min read
Watch on YouTube
The one-line take

This study finds that a carefully structured team of LLM reviewers can often critique computer architecture papers more deeply than individual human analysts, while still struggling with trust and calibration.

Key results

15
Gauntlet-preferred comparisons

Evaluators preferred Gauntlet in 15 of 20 human comparisons.

4.2
Paper 1 total-score advantage

Gauntlet's mean lead over human analysis on the 25-point rubric.

98
Ablation corpus

Papers used in the automated comparison of directive, persona, and pipeline strategies.

96%
Pipeline win rate over persona agent

Gauntlet beat the strong single-persona baseline in the Gemini 3.1 Pro ablation.

What the paper found

The paper introduces Gauntlet, an open-source pipeline for deep technical comprehension of computer-architecture papers rather than surface-level summarization. Developed by researchers from the University of Wisconsin–Madison and NVIDIA Research, Gauntlet assigns five independent reviewer agents—covering microarchitecture, workloads, simulation tools, and paper-specific specialties—to analyze each paper, then uses an adversarial synthesizer to preserve disagreements and produce a structured critique of the mechanism, key insight, evaluation, and hidden assumptions. Using Claude Opus 4.5, the study evaluated papers from ISCA 2025 and HPCA 2026 against analyses written by ten graduate researchers. Human evaluators preferred Gauntlet in 15 of 20 comparisons; its mean total-score advantage was 4.2 points in one round and 3.6 in the other on a 25-point rubric, with the strongest gains in Critical Rigor. A 98-paper ablation, judged blindly with Gemini 3.1 Pro, showed that the full pipeline beat a strong single-persona agent on 96% of papers, isolating adversarial synthesis as the main source of improvement. However, humans still won when Gauntlet made a confident technical error, explained mechanisms without teaching them, or listed weaknesses without prioritization. The authors therefore position Gauntlet as a first-pass triage and teaching aid, not a replacement for expert reading: its central contribution is the architecture of independent, domain-aware perspectives followed by disagreement-preserving synthesis.

Original abstract

Can large language models perform deep technical comprehension of computer architecture papers -- not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper through five independent expert-persona reviewers and an adversarial synthesis stage. On 20 ISCA 2025 and HPCA 2026 papers, ten researchers each wrote their own analyses and then judged, for papers other than their own, the human analysis against Gauntlet's. Across the 20 comparisons evaluators preferred Gauntlet in 15 (human in 4, one tie); its advantage is significant on per-analyst totals (paired Wilcoxon, p < 0.01) and largest on Critical Rigor, vanishing only on Calibration. Where humans win, it is on trust and usefulness rather than depth: a confident wrong claim, a mechanism described but not taught, or unprioritized breadth. A 98-paper automated ablation shows the gain comes from the multi-agent structure -- the pipeline beats the same model run as a single rich-persona agent on 96% of papers -- and specifically from its synthesis pass. We release all analyses, scores, and the rubric as a community resource.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis