NTH

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

AuthorsShicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Philip S. Yu, Lu Cheng

June 30, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a cheap, Wikipedia-based reward signal that helps language models learn to answer factual questions more accurately without relying on expensive neural verifiers.

Key results

30
models_x_benchmarks

CorVer evaluated across 30 model–benchmark cells

3
base_model_range_min

Smallest instruction-tuned base model size in billions

14
base_model_range_max

Largest instruction-tuned base model size in billions

4.1
TriviaQA_gain

Average TriviaQA improvement cited for CorVer

4.8
speedup_range_min

Lower bound of CorVer training speedup over baselines

8.4
speedup_range_max

Upper bound of CorVer training speedup over baselines

What the paper found

CorVer, from the University of Illinois Chicago, is a lightweight process-supervision method for factual QA that replaces expensive neural sentence verifiers with a corpus-grounded reward from Wikipedia co-occurrence statistics indexed by Infini-gram. It extracts a subject-object pair with a 0.5B triplet extractor, performs one corpus lookup per sentence, and maps the count into a four-tier reward that is aligned back to tokens under GRPO. Across 30 model–benchmark cells spanning six instruction-tuned models from 3B to 14B on TriviaQA, NQ-Open, PopQA, SimpleQA, and TruthfulQA, CorVer improves every Raw baseline, with the strongest average gain on Llama-3.1-8B-Instruct at +4.06 points and a +4.1-point TriviaQA gain. Against four prior factuality-RL pipelines—FoRAG, RLFH, FSPO, and KnowRL—it wins in 18 of 20 feasible comparisons while training 4.8 to 8.4× faster; average training time is 3.2 hours versus 14.5 to 29.5 hours for baselines. A human audit of 700 sentences shows monotonic calibration, rising from 24.0% correctness at co-occurrence count 0 to 81.0% at 20 or above, and ablations confirm that per-token alignment and the co-occurrence signal both matter, with collapsing the signal to a response-level scalar leaving performance below the full method.

Original abstract

Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewards provide only coarse supervision and cannot distinguish correct from incorrect statements within a reasoning trace. Sentence-level alternatives offer finer-grained feedback, but typically rely on NLI verifiers, LLM judges, or knowledge-verification pipelines that are expensive to deploy at RL scale and often unreliable for rare-entity facts, where accurate reward signals are especially important. We propose CorVer (Corpus Verify), a lightweight, plug-in-ready process reward that replaces neural verifiers with a corpus-grounded signal derived from Wikipedia co-occurrence statistics. CorVer assigns sentence-level credit and maps it to token-level advantages via a simple alignment, requiring only a 0.5B extractor and a single corpus lookup per sentence. Across 30 (model, benchmark) cells spanning six instruction-tuned models (3B to 14B) and five QA benchmarks, CorVer improves over the raw baseline for every cell, with an average TriviaQA gain of +4.1 pp. It also outperforms four neural-verifier baselines in 18 of 20 cells under their feasible configurations, while training 4.8 to 8.4x faster.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →