NTH

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

AuthorsRana Muhammad Usman

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that what an LLM agent reads in a feed can quietly steer its decisions, even when the final prompt stays the same.

Key results

4
models

Four open instruct LLMs were evaluated across the main attack grid.

10
turns

Each rollout used a ten-turn scrolling exposure phase before the decision prompt.

5
posts per turn

Each exposure turn presented five posts.

100%
Llama remote-first drop

Remote-first decisions on Llama 3.2-3B fell from 100% to 50% under heavy attack.

3×10−10
generator-swap p-value

Heavy attack with Gemma-written posts on Llama 3.2-3B.

0.85-0.95
probe accuracy

Random cross-validation accuracy for linear probes on residual-stream activations.

What the paper found

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults shows that the upstream ranker, not just the final prompt, can control what an LLM agent does after a multi-turn exposure phase. The paper tests four open instruct models—Meta’s Llama 3.2-3B, Google’s Gemma 4-e4b, and Alibaba’s Qwen 3.5-2B and Qwen 3.5-9B—using a ten-turn scrolling protocol with five posts per turn, then a forced-choice decision. On the remote-work task, heavy pro-return-to-office feeds cut Llama 3.2-3B’s remote-first rate from 100% to 50% and Gemma 4-e4b’s from 40% to 0%, while the two Qwen models stayed near hybrid defaults and did not move significantly. A generator-swap using Gemma-written posts strengthened the Llama effect further, dropping remote-first to 5% with Fisher’s exact p = 3 × 10−10. The attack showed a dose-response threshold near 2 adversarial posts per 5-post batch, generalized to security-relevant A/B/C decisions such as deployment gates and access controls, and was partially reduced by balanced exposure and ranking disclosure on Llama. The paper also reports a methodological warning: linear probes on multi-turn trajectories looked highly accurate under random cross-validation, around 0.85–0.95 balanced accuracy, but much of that signal was leaked from visible history, so group-aware evaluation is required.

Original abstract

LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts. We introduce a controlled protocol that holds the model, persona, topic, and final decision prompt fixed and varies only the composition and ordering of the posts an agent encounters during a preceding ten-turn "scrolling" phase, isolating the causal effect of feed curation on a downstream decision. Across 2,785 decision rollouts on four modern open instruct LLMs from three independent labs, we identify three response regimes: adversarial capitulation, default saturation, and a default-direction asymmetry in which a one-sided feed tips a decision the model was genuinely uncertain about (in the clearest cases from 5% to 100%; Fisher p as low as 3 x 10^-10) but cannot dislodge one it already favors or holds firmly. The effect follows a dose-response curve, survives a generator swap that rules out a writing-style artifact, generalizes across several decision domains including security-relevant choices such as removing a deployment approval gate or relaxing access controls, and is partly mitigated by two simple feed-level defenses; a frontier model retains its default. We characterize the recommender as a practical, default-bounded control surface for LLM agents, and argue that agent evaluations must audit the feed layer rather than the final prompt alone.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis