Jingyan Shen, Ang Li, Salman Rahman +4 more
Reinforcement learning has become the standard way to sharpen a language model's reasoning after the main training run. Yet the two stages are usually studied apart, which leaves two basic questions open: how do pretraining decisions like model size and data change what RL can add, and what is RL really doing to the model.
These are hard questions to answer with ordinary language models. Pretraining corpora are enormous and uncontrolled, so when a behaviour appears afterwards there is no clean way to tell whether it came from pretraining or from RL. Running systematic sweeps across both stages costs more than most researchers can spend.
So the authors use chess as a controlled setting instead. The appeal is that you get to know exactly what went into the model, which is precisely what a web-scale corpus denies you. It is the same instinct as studying genetics in fruit flies rather than in people: not because flies are the interesting case, but because control is what makes the question answerable at all.
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are…
Can We Trust Item Response Theory for AI Evaluation?
arXiv (cs.LG) · July 16, 2026RTS Smoother-Guided Learning of Physics-Based Neural Differential Models
arXiv (cs.AI) · July 16, 2026T^2MLR: Transformer with Temporal Middle-Layer Recurrence
arXiv (cs.AI) · July 16, 2026Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy