Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game
AuthorsChensong Huang, Changyu Chen, Chenwei Lin, Hanjia Lyu, Xian Xu, Jiebo Luo
Resources
This paper shows that LLMs can look like humans in risk-taking tasks while actually using very different decision mechanisms underneath.
Key results
model count in the study
original St. Petersburg game
original St. Petersburg game
models out of 28 in 20-toss truncation
models out of 28 in 20-toss truncation
mechanism transitions out of 112
What the paper found
This paper studies 28 LLMs on the St. Petersburg game to separate outcome-level resemblance from mechanism-level alignment in risk decisions. In the original unbounded game, most models appear human-like by giving finite willingness-to-pay bids, with median bids of $20 at τ = 0 and $10 at τ = 0.7, but mechanism probes expose a different pattern: under 20-toss truncation, 25 of 28 models at τ = 0 and 26 of 28 at τ = 0.7 collapse to the exact computational boundary, $21, and repeated-play, endowment, and occupational-identity variants often yield conditionally rational or boundary-tracking responses rather than human-consistent directional shifts. Human-cue prompting and instruction tuning mainly lower bids without reliably repairing the underlying response profile; for example, human-cue prompting improves only 23 of 112 mechanism transitions at τ = 0 and 19 of 112 at τ = 0.7, while instruction tuning improves 12 of 42 and 8 of 42 transitions, respectively, with most cases unchanged. An EV-first prompt that explicitly asks models to compute expected value does not eliminate the effect: finite bids still dominate, and 25 of 28 models remain computationally rational on truncation at τ = 0. The core finding is that apparently cautious LLM behavior in risk tasks can be surface-level, so evaluation should test whether decisions remain mechanism-consistent under nearby changes in horizon, wealth, and role framing, not just whether the final number looks human-like.
Original abstract
LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.