CORAL: An LLM-Native Harness for Production Recommender Systems
AuthorsMuhammad Rafay Azhar, Yuhang Zhou, Gilbert Jiang, Yuchen Wang, Rahul Sharma, Matthew DeSousa, Jiayi Liu, Xin Guo, Lizhu Zhang, Xiangjun Fan
Resources
CORAL uses an LLM agent to continuously tune live recommender systems, improving engagement or reducing serving costs while staying within operational guardrails.
Key results
Increase in video-viewing sessions across all users from fixed-budget retrieval allocation
Increase in total watch time with no additional serving cost
Increase in video-viewing sessions for new low-signal users
Increase in serving-cost savings after widening the allocation to additional segments
What the paper found
CORAL, or Constraint-Optimized Recommender via an Agentic Loop, is an LLM-native harness that lets a language-model agent continually optimize live production recommenders rather than merely rank items or develop models offline. Deployed on Meta social-platform surfaces, the system treats recommender control as a partially observed, non-stationary, constrained optimization problem: every 3-day cycle, the agent analyzes telemetry, retrieves the last 3 cycles of decisions and outcomes, estimates attribution, proposes changes, and passes them through deterministic tools, a numerical budget projection, and safety guardrails before A/B measurement feeds results back into memory. The policy improves in context without parameter updates. In a video-recommendation deployment, reallocating a fixed retrieval budget increased video-viewing sessions by 0.16% and total watch time by 0.15%, with no additional serving cost; segment-specific allocation raised sessions for new low-signal users by 0.23%. In a separate serving-capacity experiment, the agent reduced annualized compute expenditure, then expanded the intervention and increased savings by 44% while engagement remained statistically unchanged. CORAL therefore spans both sides of the engagement–efficiency frontier, compressing tuning cycles from engineer-weeks to autonomous days while retaining human oversight and enforceable operational constraints.
Original abstract
Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through online experiments--a slow, reactive process limited by engineering effort, leaving parts of the system unrevised as conditions change. Although large language models have been applied to ranking, user modeling, and offline model development, few systems place an agent in a continual closed loop that acts on a live recommender and learns from the measured effects of its decisions. We present CORAL (Constraint-Optimized Recommender via an Agentic Loop), an LLM-native harness that closes this loop: each cycle, the agent observes operating signals, reasons over a memory of past decisions and outcomes, and invokes tools--including a numerical optimizer that keeps changes within a fixed operating budget--to reconfigure the recommender, with measured outcomes informing the next cycle. We formulate this as a partially observed, non-stationary, constrained optimization problem in which the policy improves in context, without parameter updates, from its prior actions. Across two large-scale social platforms, evaluated with A/B experiments, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other, spanning the engagement-efficiency frontier. Performance improves as the loop iterates, suggesting that a single agentic loop can automate continual optimization work traditionally performed by human algorithm engineers under explicit guardrails.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.