Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
AuthorsRana Muhammad Usman
Resources
This paper shows that what an LLM agent reads in a feed can quietly steer its decisions, even when the final prompt stays the same.
Key results
Four open instruct LLMs were evaluated across the main attack grid.
Each rollout used a ten-turn scrolling exposure phase before the decision prompt.
Each exposure turn presented five posts.
Remote-first decisions on Llama 3.2-3B fell from 100% to 50% under heavy attack.
Heavy attack with Gemma-written posts on Llama 3.2-3B.
Random cross-validation accuracy for linear probes on residual-stream activations.
What the paper found
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults shows that the upstream ranker, not just the final prompt, can control what an LLM agent does after a multi-turn exposure phase. The paper tests four open instruct models—Meta’s Llama 3.2-3B, Google’s Gemma 4-e4b, and Alibaba’s Qwen 3.5-2B and Qwen 3.5-9B—using a ten-turn scrolling protocol with five posts per turn, then a forced-choice decision. On the remote-work task, heavy pro-return-to-office feeds cut Llama 3.2-3B’s remote-first rate from 100% to 50% and Gemma 4-e4b’s from 40% to 0%, while the two Qwen models stayed near hybrid defaults and did not move significantly. A generator-swap using Gemma-written posts strengthened the Llama effect further, dropping remote-first to 5% with Fisher’s exact p = 3 × 10−10. The attack showed a dose-response threshold near 2 adversarial posts per 5-post batch, generalized to security-relevant A/B/C decisions such as deployment gates and access controls, and was partially reduced by balanced exposure and ranking disclosure on Llama. The paper also reports a methodological warning: linear probes on multi-turn trajectories looked highly accurate under random cross-validation, around 0.85–0.95 balanced accuracy, but much of that signal was leaked from visible history, so group-aware evaluation is required.
Original abstract
LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts. We introduce a controlled protocol that holds the model, persona, topic, and final decision prompt fixed and varies only the composition and ordering of the posts an agent encounters during a preceding ten-turn "scrolling" phase, isolating the causal effect of feed curation on a downstream decision. Across 2,785 decision rollouts on four modern open instruct LLMs from three independent labs, we identify three response regimes: adversarial capitulation, default saturation, and a default-direction asymmetry in which a one-sided feed tips a decision the model was genuinely uncertain about (in the clearest cases from 5% to 100%; Fisher p as low as 3 x 10^-10) but cannot dislodge one it already favors or holds firmly. The effect follows a dose-response curve, survives a generator swap that rules out a writing-style artifact, generalizes across several decision domains including security-relevant choices such as removing a deployment approval gate or relaxing access controls, and is partly mitigated by two simple feed-level defenses; a frontier model retains its default. We characterize the recommender as a practical, default-bounded control surface for LLM agents, and argue that agent evaluations must audit the feed layer rather than the final prompt alone.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.