Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
AuthorsSubhadeep Pal, Shashwat Sourav, Tirthankar Ghosal, Markus J. Buehler
Resources
This paper trains AI to reason like a graph-builder, turning scientific questions into traceable chains of concepts that help generate better hypotheses for materials discovery.
Key results
Open-ended scientific questions used for evaluation
Overall improvement over corresponding base models
Relative increase in semantic diversity over baselines
Graph-PRefLexOR-8B final answers aligned with its own reasoning stages in 92 of 100 cases
Nodes in the final leap graph after test-time expansion
Edges in the final leap graph after test-time expansion
What the paper found
Graph-PRefLexOR is a graph-native reinforcement learning framework for scientific hypothesis generation that turns the hidden reasoning trace into an explicit sequence of <brainstorm>, <graph>, <graph_json>, <patterns>, and <synthesis> stages, then optimizes that structure with ORPO cold start and Group Relative Policy Optimization (GRPO). Built on Qwen3-1.7B, Llama-3.2-3B-Instruct, and Qwen3-8B backbones, the system is evaluated on 100 open-ended materials science and mechanics questions judged by Claude Opus-4.7 on 0–10 scores, where it improves aggregate reasoning performance by 40–65 percent over the corresponding base models, with the largest gains in traceability. Embedding analysis with google/embeddinggemma_300m and BAAI/bge-base-en-v1.5 shows roughly 2×–3× greater semantic diversity, including inter-phase cosine-distance gains from 0.07 to 0.20 at 1.7B and from 0.08 to 0.21 at 8B. Semantic backtracking shows Graph-PRefLexOR-8B final answers align with its own structured reasoning in 92 of 100 cases, and specifically with <synthesis> in 89 of 100 cases, while Qwen3-8B aligns with its own thinking trace in only 16 of 100 cases. Test-time graph expansion over 2,000 iterations on self-healing biopolymer composites shows that explored semantic volume saturates within a few hundred iterations, but surprising recombinations continue to grow super-linearly, culminating in a leap graph with 4,419 nodes and 37,064 edges and revealing that additional compute mainly densifies a bounded idea space rather than expanding it indefinitely.
Original abstract
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph construction, pattern extraction, and hypothesis synthesis. This design links neural language generation with symbolic relational structure, enabling causal connections to be constructed, inspected, and reused. On 100 open-ended questions from materials science and mechanics literature, Graph-PRefLexOR achieves 40-65% improvements over corresponding base models, with the largest gains in reasoning traceability. Embedding analyses show broader semantic exploration and approximately 2-3 times greater semantic diversity than baselines. Semantic backtracking and layer-wise hidden-state analyses further show stronger alignment between structured reasoning and final answers. Finally, test-time graph expansion reveals that additional compute primarily increases long-range conceptual recombination within a bounded semantic space, rather than simply expanding semantic coverage. These results establish graph-native reinforcement learning as a pathway toward interpretable AI systems for scientific hypothesis generation in materials design and other scientific applications.
Read the original paperMore in Graph Learning
Browse all 32 papers →CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
CodeGraph turns 167 million source files into a Wikidata-grounded knowledge graph of algorithms, paradigms, patterns, and software domains.
GraphWrit3R: End-to-End 3D Scene Graph Writing
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.
Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
This work shows how a small, carefully designed set of single-variable interventions can recover causal structure even when hidden confounders and selection bias complicate the data.