Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
AuthorsParand A. Alamdari, Toryn Q. Klassen, Sheila A. McIlraith
Resources
This paper brings formal verification ideas to LLM safety, showing how to audit, monitor, and even intervene at runtime to keep AI systems within rules and constraints.
Key results
TRACP+I used predictive monitoring over a 3-step horizon before triggering interventions.
What the paper found
This paper argues that AI governance needs product-level compliance monitoring, not just model-level safety. The core contribution is TRAC, a black-box framework that audits and monitors advanced AI systems, especially LLM agents, against temporally extended constraints encoded in Linear Temporal Logic (LTL). Instead of asking an LLM judge to reason over entire behavior logs, TRAC decomposes the problem into a labeling function L that extracts atomic propositions from each input-output step and an LTL progression operator that symbolically updates the residual formula online, producing three-valued verdicts and interpretable execution witnesses. In three environments—IPC-Trucks compiled from PDDL3 preferences, TextWorld cooking, and ScienceWorld—TRACR with even small LLM labelers outperformed direct LLM-as-a-Judge baselines on violation detection; the paper reports that small labeling models within TRACR matched or exceeded frontier judges such as Gemini 2.5 Pro, while LLM judges degraded sharply as event distance, constraint count, and proposition count increased. The paper also introduces TRACP+I, which adds predictive monitoring over a 3-step horizon and black-box interventions via best-of-n resampling, constraint-guided prompting, or safer-model substitution; across all tested model-environment pairs it reduced violation rates without significant task-performance loss. The key novelty is the formal-methods/LLM split: LLMs handle local semantic grounding, while LTL progression handles temporal correctness, yielding more reliable auditing and proactive control than prompting-only approaches.
Original abstract
We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle, from pre-deployment testing to post-deployment auditing. Combining principles from formal methods with SoTA machine learning, we propose techniques that enable AI-enabled product and service developers, as well as third party AI developers and evaluators, to perform offline auditing and online (runtime) monitoring of product-specific (temporally extended) behavioral constraints such as safety constraints, norms, rules and regulations with respect to black-box advanced AI systems, notably LLMs. We further provide practical techniques for predictive monitoring, such as sampling-based methods, and we introduce intervening monitors that act at runtime to preempt and potentially mitigate predicted violations. Experimental results show that by exploiting the formal syntax and semantics of Linear Temporal Logic (LTL), our proposed auditing and monitoring techniques are superior to LLM baseline methods in detecting violations of temporally extended behavioral constraints; with our approach, even small-model labelers match or exceed frontier LLM judges. Our predictive and intervening monitors significantly reduce the violation rates of LLM-based agents while largely preserving task performance. We further show through controlled experiments that LLMs' temporal reasoning shows a pronounced degradation in accuracy with increasing event distance, number of constraints, and number of propositions.
Read the original paper