NTH

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

AuthorsOmatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho

August 12, 2026 2 min read
Watch on YouTube
The one-line take

Instead of trusting the most common LLM answer, CALVER checks whether each causal explanation actually obeys the rules of causal graphs.

Key results

42.1%
CLEAR CALVER accuracy

Selection accuracy on 126 multi-answer CLEAR problems at K = 8.

30.5%
Skywork Reward V2 8B accuracy

Best generic reward-model comparator on the same frozen candidate pools.

57.9%
Qwen2.5-7B SFT CALVER at K = 32

Accuracy after scaling from K = 1 to K = 32.

32.5%
Plurality at K = 32

Exact-plurality accuracy on the corresponding frozen Qwen2.5-7B SFT pools.

7%
Inference overhead

Approximate additional cost of symbolic verification at K = 8.

What the paper found

The paper introduces CALVER, a training-free symbolic verifier for best-of-K causal reasoning in large language models, addressing a failure mode of self-consistency: several answers may satisfy the same causal predicate, so voting fragments correct probability mass while an invalid answer becomes the most frequent string. CALVER requires each sampled trace to use six typed slots, then deterministically checks graph binding, query binding, d-separation or m-separation, backdoor-adjustment validity, derivation provenance, numerical recomputation, and answer consistency. On 126 multi-answer problems from the CLEAR benchmark, CALVER achieved 42.1% selection accuracy at K = 8, versus 30.5% for Skywork Reward V2 8B and near 30% for exact plurality and a reference-free language-model judge; scaling that judge to Qwen2.5-72B still did not close the gap. With Qwen2.5-7B SFT, CALVER increased accuracy from 20.6% at K = 1 to 57.9% at K = 32, while plurality plateaued at 32.5%. The method transferred across ten bnlearn Bayesian networks, graph structures reconstructed from prose, Mistral NeMo 12B, and Knights-and-Knaves logic checked by truth tables. Verification added about 7% to sampling cost, with each trace scored in 1–8 milliseconds on CPU. The study also identifies a crossover: extract-then-solve is better when one graph can be recovered reliably, while candidate-wise symbolic verification is preferable when interpretations diverge and multiple valid answers preserve different query-relevant relations.

Original abstract

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. On CLEAR find-one-valid queries that admit multiple graph-valid answers, CALVER reaches 42.1% where plurality, a reward model, an LLM judge, and model confidence remain near 30% on identical frozen pools. Scaling the judge to 72B does not close the gap. In an audited clean-core subset, 11 of 21 graph-valid CALVER selections differ from the benchmark's listed answer while still satisfying the requested predicate. The advantage widens with the sampling budget and reproduces across ten published Bayesian networks, a second model family, and settings where the model must build the graph from text. CALVER also improves thresholded average-treatment-effect decisions against exact ground truth, generalizes to logic under a truth-table checker, and scores each candidate in milliseconds on CPU. CALVER needs only a causal structure, supplied outright or built from the text; wherever that holds, selection can aggregate via causal validity.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis