CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
AuthorsFederico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
AffiliationsUniversità di Bologna, Bologna, Italy · LTCI, Télécom Paris, Institut Polytechnique de Paris, Palaiseau, France · DISI, Università di Bologna, Bologna, Italy · Sant’Anna School of Advanced Studies, Pisa, Italy
Resources
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
Key results
Source files in the corpus processed by the pipeline.
Total nodes in the resulting typed knowledge graph.
Typed edges linking files, concepts, Wikidata entities, and hierarchy relations.
Distinct Wikidata entities imported during grounding.
Devstral-2 2512 accuracy against the human gold standard.
Devstral-2 2512 F1 score against the human gold standard.
What the paper found
CodeGraph introduces an open-taxonomy knowledge-graph pipeline that converts source code into semantic annotations for application domains, algorithms, programming paradigms, and design patterns. Applied to the 167M-file Stack-Edu corpus, it retains 145M executable-source files across 14 languages and produces 158M graph nodes connected by 1.02B typed edges. The extraction stage uses Qwen3-Coder-30B-A3B-Instruct, a code-specialized mixture-of-experts model, while grounding proceeds through deterministic Wikidata SPARQL matching, Qwen3.6-27B Deep Research Agent disambiguation for ambiguous labels, and parent-hierarchy rollup. The resulting graph contains 19,807 grounded Wikidata entities, enabling semantic code search and cross-domain algorithm discovery beyond lexical or token-level retrieval. Quality assurance combines a human gold set with LLM verification: among Claude Haiku 4.5, Gemini 3 Flash, GPT-5.1-Codex-Mini, Grok Code Fast 1, Devstral-2 2512, and DeepSeek V3.2, Devstral-2 2512 performed best at 80.6% accuracy, 82.1% precision, 96.0% recall, and 88.5% F1. On a 10,000-file silver-tier sample, verifier acceptance reached 93% for paradigms, 88% for algorithms and domains, and 82% for design patterns, although the authors caution that design-pattern and paradigm labels are taxonomically less reliable. Construction required approximately 400,000 GPU-hours on NVIDIA A100 hardware, making CodeGraph a large-scale semantic layer for source-code analysis rather than merely another code corpus.
Original abstract
Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.
Read the original paperMore in Graph Learning
Browse all 32 papers →GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.
Efficient Swing Computation for Retrieval in Large-Scale Recommender Systems
Runhao Jiang, Renchi Yang
This paper makes massive-scale item recommendations much faster by approximating graph-based similarity scores without sacrificing retrieval quality.