Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
AuthorsShicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Philip S. Yu, Lu Cheng
Resources
This paper introduces a cheap, Wikipedia-based reward signal that helps language models learn to answer factual questions more accurately without relying on expensive neural verifiers.
Key results
CorVer evaluated across 30 model–benchmark cells
Smallest instruction-tuned base model size in billions
Largest instruction-tuned base model size in billions
Average TriviaQA improvement cited for CorVer
Lower bound of CorVer training speedup over baselines
Upper bound of CorVer training speedup over baselines
What the paper found
CorVer, from the University of Illinois Chicago, is a lightweight process-supervision method for factual QA that replaces expensive neural sentence verifiers with a corpus-grounded reward from Wikipedia co-occurrence statistics indexed by Infini-gram. It extracts a subject-object pair with a 0.5B triplet extractor, performs one corpus lookup per sentence, and maps the count into a four-tier reward that is aligned back to tokens under GRPO. Across 30 model–benchmark cells spanning six instruction-tuned models from 3B to 14B on TriviaQA, NQ-Open, PopQA, SimpleQA, and TruthfulQA, CorVer improves every Raw baseline, with the strongest average gain on Llama-3.1-8B-Instruct at +4.06 points and a +4.1-point TriviaQA gain. Against four prior factuality-RL pipelines—FoRAG, RLFH, FSPO, and KnowRL—it wins in 18 of 20 feasible comparisons while training 4.8 to 8.4× faster; average training time is 3.2 hours versus 14.5 to 29.5 hours for baselines. A human audit of 700 sentences shows monotonic calibration, rising from 24.0% correctness at co-occurrence count 0 to 81.0% at 20 or above, and ablations confirm that per-token alignment and the co-occurrence signal both matter, with collapsing the signal to a response-level scalar leaving performance below the full method.
Original abstract
Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewards provide only coarse supervision and cannot distinguish correct from incorrect statements within a reasoning trace. Sentence-level alternatives offer finer-grained feedback, but typically rely on NLI verifiers, LLM judges, or knowledge-verification pipelines that are expensive to deploy at RL scale and often unreliable for rare-entity facts, where accurate reward signals are especially important. We propose CorVer (Corpus Verify), a lightweight, plug-in-ready process reward that replaces neural verifiers with a corpus-grounded signal derived from Wikipedia co-occurrence statistics. CorVer assigns sentence-level credit and maps it to token-level advantages via a simple alignment, requiring only a 0.5B extractor and a single corpus lookup per sentence. Across 30 (model, benchmark) cells spanning six instruction-tuned models (3B to 14B) and five QA benchmarks, CorVer improves over the raw baseline for every cell, with an average TriviaQA gain of +4.1 pp. It also outperforms four neural-verifier baselines in 18 of 20 cells under their feasible configurations, while training 4.8 to 8.4x faster.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.