NTH
AI research

Who Owns This Agent? Tracing AI Agents Back to Their Owners

AuthorsRuben Chocron, Doron Jonathan Ben Chayim, Eyal Lenga, Gilad Gressel, Alina Oprea, Yisroel Mirsky

May 18, 2026 3 min read
Watch on YouTube
The one-line take

This paper proposes a way for AI vendors to trace a harmful autonomous agent back to the account that launched it, helping close a major accountability gap.

Key results

TPR ≈ 1.0 with FPR ≈ 0.0
Communication agents lexical canaries

In chat and forum experiments, simple lexical canaries were nearly perfectly recovered in non-adversarial settings.

macro AUC 0.98
Chat semantic canary detector AUC

A multilabel GRU probe over frozen Qwen 2.5-7B-Instruct activations detected semantic canaries in chat.

TPR ≈ 1.00 at FPR ≈ 0.01 with 20 semantic canaries
Chat semantic canary session attribution

Session-level attribution on chat remained highly accurate when using an m-of-k rule over 20 semantic canaries.

average AUC 0.992 non-adversarial and 0.974 after paraphrase
Web semantic canary detector AUC

A ModernBERT-based multilabel classifier recovered semantic canaries from HTML both before and after paraphrase.

51% to 9%
Web-CTF solve rate drop

With a canary-aware adversary, utility-bearing canaries reduced agent solve rate in 15 Flask-based web-CTF challenges.

about 2× faster
Search speedup

Vendor-side token-level matching was about 2× faster than character-level search over session logs.

What the paper found

Who Owns This Agent? introduces agent attribution, the problem of tracing an observed AI agent interaction back to the operator’s vendor account. The paper’s core contribution is a canary-based protocol for vendor-hosted LLMs: an authorized authority injects a canary into content the agent is likely to consume, and the vendor searches a bounded session-log window around the injection time to recover the session and account. In non-adversarial settings, a single lexical canary is enough and achieves near-perfect attribution; in communication-agent experiments on Qwen 2.5-7B-Instruct chat and forum data, lexical canaries reached TPR ≈ 1.0 with FPR ≈ 0.0. For adversarial operators who paraphrase input, the authors propose utility-bearing lexical canaries and semantic canaries, where meaning rather than exact token sequence is detected by classifiers over cached representations. On chat, a multilabel GRU probe over frozen Qwen 2.5-7B-Instruct activations achieved macro AUC 0.98 and session-level TPR ≈ 1.00 at FPR ≈ 0.01 with 20 semantic canaries; on web pages, a ModernBERT-based multilabel classifier reached average AUC 0.992 non-adversarial and 0.974 after paraphrase. In 15 Flask-based web-CTF challenges using Qwen 3.5-9B, utility-bearing canaries showed the key asymmetry: a canary-aware adversary cut solve rate from 51% to 9% by stripping task-critical content, while attribution stayed above 90% at low false-positive rates. Vendor-side search scaled linearly with the number of sessions in the attribution window and token-level matching was about 2× faster than character search.

Original abstract

AI agents are increasingly deployed to act autonomously in the world, yet there is still no reliable way to trace a harmful agent back to the account that deployed it. This creates the same accountability gap across both ends of the intent spectrum: benign operators may deploy misconfigured or overbroad agents that cause harm unintentionally, while malicious operators may deliberately weaponize agents for scams, harassment, or cyber attacks. In many cases, these agents are powered by vendor-hosted models, a dependency that holds even for sophisticated adversaries such as state actors conducting cyber operations. In either case, affected parties can observe the behavior but cannot notify the responsible operator, stop the session, or identify the account for investigation. We formalize this gap as the problem of agent attribution: linking an observed agent interaction to the responsible account at the hosting vendor. To our knowledge, this is the first work to define the problem and present a practical solution. Our protocol is canary-based: an authorized party injects a canary into the agent's interaction stream, and the vendor searches a narrow window of session logs to recover the originating session and account. Simple canaries suffice in non-adversarial settings. For adversarial operators who filter or paraphrase incoming content, we develop robust canary constructions that cannot be suppressed without degrading the agent's own task performance, yielding a formal asymmetry in the defender's favor. We evaluate a variety of scenarios including real-world agents and show that our attribution method is reliable, robust, and scalable for vendor-side deployment.

Read the original paper