NTH
AI research

Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis

AuthorsPenghui Li, Hong Yau Chong, Yinzhi Cao, Junfeng Yang

May 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces Neo, an AI-assisted program analysis system that hunts for privilege-escalation bugs across large polyglot microservice codebases and has already found dozens of real-world vulnerabilities.

Key results

25 open-source microservice applications
Evaluation corpus

NEO was evaluated across 25 real-world microservice apps spanning 7 languages and 6.2 million lines of code.

20 verified vulnerabilities across 4 applications
Ground-truth dataset

The paper’s ground-truth set includes 18 previously reported vulnerabilities plus 2 newly uncovered ones.

24
Zero-day privilege escalation vulnerabilities

NEO uncovered 24 previously unknown privilege escalation vulnerabilities in the evaluation.

81.0%
Ground-truth precision

On the ground-truth dataset, NEO achieved 81.0% precision for privilege escalation detection.

85.0%
Ground-truth recall

On the ground-truth dataset, NEO achieved 85.0% recall for privilege escalation detection.

76.5%
Overall precision

Across both datasets, NEO reported 39 true positives and 12 false positives, yielding 76.5% precision overall.

What the paper found

Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis presents NEO, an LLM-driven security analysis framework that combines Claude Sonnet 3.7 with CodeQL-based program analysis to detect cross-service privilege escalation in heterogeneous microservice codebases. The core novelty is a small set of language-agnostic code search primitives—Qname, Qast, Qflow, Qcg, Qsource, Qinter, and Qglobalflow—that let the agent iteratively identify privileged operations, trace data from external inputs across service boundaries, and validate whether authN/authZ checks actually protect a specific sink. Evaluated on 25 open-source microservice applications spanning 7 languages and 6.2 million lines of code, NEO uncovered 24 zero-day privilege escalation vulnerabilities and achieved 81.0% precision and 85.0% recall on a ground-truth set of 20 verified vulnerabilities across 4 applications; across both datasets it reported 39 true positives with 12 false positives, for 76.5% precision overall. The ablation study shows why the design matters: removing application-specific privileged operation discovery drops recall from 92.9% to 33.3%, and removing the code-search primitives reduces recall to 4.8%. Compared with MScan, CodeQL, and EnIGMA, NEO found substantially more vulnerabilities, including 24 more than EnIGMA in one comparison set. Beyond privilege escalation, prompt-only retargeting let it find 18 additional zero-day command-injection and SQL-injection issues, showing that the agentic search-and-validate loop generalizes beyond a single vulnerability class.

Original abstract

Microservices are widely adopted in modern cloud systems due to their scalability and fault tolerance. However, microservice architectures introduce significant complexity in privilege and permission control, creating risks of privilege escalation where attackers can gain unauthorized access to resources or operations. Detecting such vulnerabilities is challenging due to complex cross-service interactions, polyglot codebases, and diverse privileged operations and permission checks. We present Neo, an agentic program analysis framework that combines large language models (LLMs) with classic program analysis to address these challenges. Neo leverages an LLM-based agent that dynamically generates analysis plans, adapts code search strategies, and validates semantics. We develop code search primitives that enable Neo to perform scalable and flexible code exploration across services and languages. We evaluated Neo on 25 open-source microservice applications spanning 7 programming languages and 6.2 million lines of code. Neo uncovered 24 zero-day privilege escalation vulnerabilities and achieved 81.0% precision and 85.0% recall on a ground-truth dataset. Compared to existing program analysis and agentic solutions, Neo demonstrated significant improvements in both detection accuracy and scalability. We further showcased Neo's extensibility by applying it to other application domains and vulnerability types, uncovering 18 additional zero-day vulnerabilities.

Read the original paper