Learning to Trigger: Reinforcement Learning at the Large Hadron Collider
AuthorsZixin Ding, Shaghayegh Emami, Giovanna Salvi, Cecilia Tosciri, Abhijith Gandrakota, Jennifer Ngadiuba, Nhan Tran, Christian Herwig, David W. Miller, Yuxin Chen
This paper uses reinforcement learning to dynamically tune particle-collider trigger thresholds, improving real-time event selection on both simulated and real LHC data.
Key results
fraction of chunks within tolerance on the HT trigger
fraction of chunks within tolerance on the AD trigger
zero-shot transfer on CMS Run 283408 for the HT trigger
zero-shot transfer on CMS Run 283408 for the AD trigger
best F1 on the Numenta Anomaly Benchmark
What the paper found
This paper from the University of Chicago, the University of Michigan, and Fermilab reframes Large Hadron Collider trigger tuning as streaming reinforcement learning, where a policy updates a single threshold online to hold background rate inside a tight tolerance band while maximizing signal efficiency. Using CMS Run 283408 as the real-data deployment target, the authors show that a basic DQN already improves over static menus, but critic-based and Lagrangian methods break under drift: GRPO and L-GRPO often encounter zero-feasible candidate groups, and CPO becomes effectively degenerate. Their main contribution is Group-Filtered Policy Optimization, or GFPO, with two variants: GFPO-F filters candidate threshold updates by smallest rate error, while GFPO-FR first keeps feasible candidates and then ranks them by signal utility. On Monte Carlo streams, GFPO-F raises in-band fraction to 1.000 on the HT trigger and 1.000 on the AD trigger, while GFPO-FR reaches 1.000 on HT and 0.979 on AD; on CMS data without fine-tuning, GFPO-FR achieves 0.950 on HT and 0.608 on AD, and GFPO-F reaches 0.990 on HT and 0.689 on AD. The paper also introduces a sequence-based state encoder with a GRU and a 22-feature observation vector, and shows the approach transfers beyond particle physics: on UNSW-NB15 it preserves false-alert-rate control, and on NAB it reaches 0.215 to 0.216 F1, beating an oracle static threshold at 0.184.
Original abstract
High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48\% ($H_T$) and 28\% (AD), with a cumulative gain of up to 2\% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56\% ($H_T$) and 28\% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO_LHC (see repo for details).
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.