SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
AuthorsShuofei Qiao, Yunxiang Wei, Jiazheng Fan, Bin Wu, Busheng Zhang, Mengru Wang, Yuqi Zhu, Ningyu Zhang, Keyan Ding, Qiang Zhang, Huajun Chen
Resources
SciAtlas builds a massive cross-disciplinary knowledge graph for science and pairs it with graph-based retrieval to help AI agents do better literature search, synthesis, and research planning.
Key results
SciAtlas covers 43.30M academic papers
Total entity count in SciAtlas
Total relational triplets in SciAtlas
SciAtlas spans 26 academic disciplines
Paper and keyword vector indexes use 1024-dimensional embeddings
What the paper found
SciAtlas is a large-scale, multi-disciplinary knowledge graph for automated scientific research that targets the failure modes of keyword search, vector retrieval, and LLM deep-research workflows, especially their weak topological reasoning and hallucination risk. Built from OpenAlex and deployed in Neo4j, it organizes 43.30M papers across 26 disciplines into 157M entities and 3B triplets, with 9 node types, 12 relation types, and 1024-dimensional embeddings over titles, abstracts, and keywords. A lightweight Qwen3-30B-A3B-Instruct-2507 extractor generates 3–8 canonical keywords per paper, while bge-large-en-v1.5 and bge-reranker-large support hybrid semantic retrieval. The core retrieval pipeline uses tri-path collaborative recall—keyword matching, semantic matching, and title matching—followed by 2-hop graph propagation with random walk with restart, citation-aware weighting, and graph reranking; it returns the top-20 papers with path-based explanations in under 2 minutes. The paper’s novelty is not a benchmark jump but a deterministic, neuro-symbolic “cognitive map” for literature review, idea grounding, trend synthesis, author retrieval, and researcher profiling, designed to reduce inference cost while preserving deep association discovery.
Original abstract
The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowledge organization impedes deep interdisciplinary integration. Current academic retrieval tools predominantly rely on superficial keyword matching or vector-space semantic retrieval, which lack the topological reasoning capabilities required to navigate complex logical connections. Agentic deep-research-based frameworks are often prone to logical hallucinations and consuming high inference costs. To bridge this gap, in this report, we introduce SciAtlas, a large-scale, multi-disciplinary, heterogeneous academic resource knowledge graph designed as a panoramic scientific evolution network. By integrating over 43M papers from 26 disciplines, and a total of 157M entities and 3B triplets, SciAtlas provides a structured topological cognitive substrate that dismantles disciplinary barriers and furnishes AI agents with a global perspective. Furthermore, we develop a neuro-symbolic retrieval algorithm featuring tri-path collaborative recall and graph reranking, achieving a seamless transition from simple semantic matching to deterministic association discovery. We also present key application directions of SciAtlas, including literature review, automated research trend synthesis, idea positioning, and academic trajectory exploration, to demonstrate that SciAtlas can serve as an effective ``cognitive map'' to empower the full loop of automated scientific research while significantly reducing reasoning costs. We have released the interfaces for KG retrieval and various downstream tasks in our GitHub repo.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.