STINER: Automated Extraction of Strategic Cyber Threat Intelligence from X
AuthorsYasir Ech-Chammakhy, Oussama Azrara, Jaafar Chbili, Anas Motii
Resources
STINER turns noisy X posts into structured cyber-threat intelligence, showing that domain-specific models can extract strategic signals faster and more accurately than general-purpose or LLM-based systems.
Key results
Expert-annotated X alerts in the STINER corpus
Best exact span-level extraction score
Inference time per tweet
Strict F1 improvement from 27.78% to 74.55%
Estimated gigabytes of leaked data in H1 2025
What the paper found
STINER is a framework for extracting strategic cyber threat intelligence from X, formerly Twitter, where ransomware claims and breach disclosures can appear before formal reporting. Its expert-annotated corpus contains 2100 real-world alerts labeled with eight entity types: TARGET, ACTOR, SECTOR, LOCATION, SIZE, DATA_TYPE, PRICE, and DATE. The benchmark evaluates nine models across 12 configurations, including BERT-Base, Twitter-RoBERTa, DarkBERT, GLiNER-Large-v2.1, and generative models Llama-3, Gemma-2, and Qwen-2.5, using strict span-level F1 on a chronological holdout set. DarkBERT achieves the best score at 89.33% and processes each tweet in 0.71 ms, outperforming fine-tuned Llama-3, Gemma-2, and Qwen-2.5 models that use QLoRA. QLoRA raises Llama-3 from 27.78% to 74.55% strict F1, a 47-point gain, but generative extraction remains slower and less accurate. The resulting STINER-DarkBERT system reconstructs European threat activity for H1 2025, identifies Akira and Qilin as leading threats in Spain, and detects a SafePay activity surge approximately 6–7 weeks before its consolidation in quarterly reporting. Aggregated extraction also estimates 146018 GB of leaked manufacturing data, showing how social-media mining can quantify strategic impact rather than merely count incidents.
Original abstract
Strategic Cyber Threat Intelligence (CTI) focuses on high-level insights, such as identifying targeted industries, attributing attacks to specific ransomware groups, and assessing the scale of data loss. Today, X (formerly Twitter) has become the fastest source for this intelligence, often hosting real-time breach announcements days before formal vendor reports. Converting this raw chatter into actionable intelligence requires navigating a complex linguistic landscape. Conventional Named Entity Recognition (NER) models struggle to parse the informal and highly irregular dialect of social media, creating a blind spot for automated defense systems. To address this challenge, we introduce STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams. We construct a high-quality, expert-annotated dataset of 2,100 real-world alerts and propose a granular taxonomy of eight entity types centered on strategic pivots such as Threat Actor, Sector, and Location. We benchmark nine models across 12 evaluated configurations, spanning general-purpose and domain-adapted encoders, open-schema extraction, and generative LLMs in both zero-shot and fine-tuned settings. Domain-adapted encoders such as DarkBERT reach a strict F1-score of 89.33%, outperforming both general-purpose baselines and fine-tuned Large Language Models, which additionally incur substantially higher inference latency. Leveraging STINER-DarkBERT, we conduct a European threat landscape analysis for H1 2025. Our results align with official reporting on major targets while highlighting the distinct visibility profile of attacks in Spain, and illustrate how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.