Kai Ruan, Jinghao Lin, Zihe Huang +4 more
Muon holds its own against AdamW when pre-training at scale. Whether that advantage carries into reinforcement learning post-training is a separate question, and this study tests it in the specific setting of sparse-reward agentic RL on ALFWorld with a small Qwen model.
The interesting result is conditional rather than general. Applying Muon only to hidden weight matrices lifted final-window validation success substantially, which is a finding about where the optimiser helps rather than whether it helps. Different parameter groups apparently want different treatment.
The methodology deserves a note too: matched single-seed comparisons are honest about their own limits. One seed cannot separate a real effect from variance, and reporting it plainly is better than the alternative common in this literature, which is to run several and present the flattering one.
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no…
Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers
arXiv (cs.LG) · July 17, 2026Pick-to-Learn Calibration of an MPC Policy for an Origin-to-Destination Flight Problem
arXiv (cs.LG) · July 17, 2026Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting
arXiv (cs.CV) · July 17, 2026Adaptive Contrast Enhancement and Optimised Feature Matching for RootSIFT-Based Palm-Vein Recognition
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.