When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
AuthorsOmatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho
Resources
Instead of trusting the most common LLM answer, CALVER checks whether each causal explanation actually obeys the rules of causal graphs.
Key results
Selection accuracy on 126 multi-answer CLEAR problems at K = 8.
Best generic reward-model comparator on the same frozen candidate pools.
Accuracy after scaling from K = 1 to K = 32.
Exact-plurality accuracy on the corresponding frozen Qwen2.5-7B SFT pools.
Approximate additional cost of symbolic verification at K = 8.
What the paper found
The paper introduces CALVER, a training-free symbolic verifier for best-of-K causal reasoning in large language models, addressing a failure mode of self-consistency: several answers may satisfy the same causal predicate, so voting fragments correct probability mass while an invalid answer becomes the most frequent string. CALVER requires each sampled trace to use six typed slots, then deterministically checks graph binding, query binding, d-separation or m-separation, backdoor-adjustment validity, derivation provenance, numerical recomputation, and answer consistency. On 126 multi-answer problems from the CLEAR benchmark, CALVER achieved 42.1% selection accuracy at K = 8, versus 30.5% for Skywork Reward V2 8B and near 30% for exact plurality and a reference-free language-model judge; scaling that judge to Qwen2.5-72B still did not close the gap. With Qwen2.5-7B SFT, CALVER increased accuracy from 20.6% at K = 1 to 57.9% at K = 32, while plurality plateaued at 32.5%. The method transferred across ten bnlearn Bayesian networks, graph structures reconstructed from prose, Mistral NeMo 12B, and Knights-and-Knaves logic checked by truth tables. Verification added about 7% to sampling cost, with each trace scored in 1–8 milliseconds on CPU. The study also identifies a crossover: extract-then-solve is better when one graph can be recovered reliably, while candidate-wise symbolic verification is preferable when interpretations diverge and multiple valid answers preserve different query-relevant relations.
Original abstract
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. On CLEAR find-one-valid queries that admit multiple graph-valid answers, CALVER reaches 42.1% where plurality, a reward model, an LLM judge, and model confidence remain near 30% on identical frozen pools. Scaling the judge to 72B does not close the gap. In an audited clean-core subset, 11 of 21 graph-valid CALVER selections differ from the benchmark's listed answer while still satisfying the requested predicate. The advantage widens with the sampling budget and reproduces across ten published Bayesian networks, a second model family, and settings where the model must build the graph from text. CALVER also improves thresholded average-treatment-effect decisions against exact ground truth, generalizes to logic under a truth-table checker, and scores each candidate in milliseconds on CPU. CALVER needs only a causal structure, supplied outright or built from the text; wherever that holds, selection can aggregate via causal validity.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.