FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection
AuthorsHuanchi Wang, Zihang Huang, Yifang Tian, Kristina Dzeparoska, Hans-Arno Jacobsen, Alberto Leon-Garcia
Resources
FAME uses an LLM only once offline to help build a lightweight on-premise expert system that pinpoints anomalous log messages with far less labeling effort.
Key results
At K = 100 on the Blue Gene/L benchmark, FAME achieves message-level anomaly detection F1 of 98.16.
On BGL at K = 100, the model reports a high-precision, high-recall operating point supporting the 98.16 F1 result.
At K = 100 on BGL, this is the human labeling budget used after K-shot sampling per EventID.
The paper states that 53,287 labeled lines correspond to a 76× reduction versus full line-level annotation on BGL.
FAME detects 86.3% of anomalies from EventIDs not seen during setup on BGL, showing generalization to unseen templates.
On Thunderbird, FAME reaches near-perfect message-level anomaly detection with F1 of 99.95 and perfect recall.
What the paper found
FAME, or Failure-Aware Mixture-of-Experts, reframes log anomaly detection from window-level scoring to message-level classification by combining one offline LLM-driven semantic partition with fully on-premise inference. The method parses raw logs with Drain3, samples at most K labeled lines per EventID, then uses an LLM once to group templates into failure domains; a deterministic certification step keeps only partitions consistent with K-shot evidence. FAME exploits an asymmetric property of rare-anomaly sampling: if all K labels in a domain are anomalous, routing alone is statistically sufficient, while all-normal samples still require a classifier. This enables a sparse MoE with a DistilBERT gate, a multiclass selector, and per-domain BERT experts, with a special UNIVERSAL_NORMAL fallback expert calibrated to catch misrouted anomalies. On BGL, the hardest benchmark because normal and anomalous lines often share templates, FAME reaches F1 = 98.16 at K = 100 with 98.18 precision and 98.14 recall, using 53,287 labeled lines, a 76× reduction versus full line-level annotation; it also detects 86.3% of anomalies from unseen EventIDs. On Thunderbird, where anomalies are structurally separable, it achieves F1 = 99.95 with perfect recall. Compared with direct in-context LLM classification, which reaches up to F1 = 96.92 on BGL and incurs $6,698–$10,047 per run for frontier models, FAME costs $10.23 one-time and processes up to 1.20 million lines per hour on a single 4-GPU node. Ablations show that a single global BERT collapses to F1 = 29.65, while removing LLM grouping still leaves a strong TF-IDF variant at F1 = 93.01, indicating that routed specialization is the main source of gain.
Original abstract
Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible. This coarse granularity forces operators to inspect many routine lines per alert. Message-level detection offers finer granularity, but remains challenging. A single event template may correspond to both normal and anomalous messages, failures arise from heterogeneous subsystems, and line-level labeling at scale is impractical. Although large language models (LLMs) can reason over log semantics, applying them to every line is too costly for continuous monitoring. We present FAME (Failure-Aware Mixture-of-Experts), a label-efficient message-level mixture-of-experts framework that uses an LLM only once offline. We annotate at most K labeled lines per template to derive binary normal/anomaly indicators and representative examples. The LLM proposes a partition of templates into failure domains, and a certification step validates the proposal before training. FAME trains a lightweight router and domain experts that run on-premise and output anomaly predictions and failure-domain labels. On BGL, FAME achieves F1 = 98.16 at K = 100 reducing annotation effort by 76x and detects 86.3% of anomalies from unseen EventIDs. On Thunderbird, FAME reaches F1 = 99.95 with perfect recall.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.