S. Aaron McClendon
Model merging is offered as a substitute for joint multi-task training, and the authors notice something awkward about the evidence. Merging is normally applied to independently released agents precisely because no joint model exists, which means the comparison against the thing it claims to replace is almost never run.
So they run it. Two Qwen3-8B specialists trained on different difficulty levels of the AppWorld agent benchmark with LOOP, merged using TIES and RAM+, and measured against the joint model that merging was supposed to stand in for.
Building a missing baseline is unglamorous and disproportionately valuable. A whole line of work can accumulate confidence from comparisons among its own variants, and the only way to find out whether the premise holds is to construct the control everyone had a practical reason to skip.
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result…
Cluster-Aware Matching via Laplacian Optimal Transport
arXiv (cs.LG) · July 17, 2026Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
arXiv (cs.AI) · July 17, 2026Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
arXiv (cs.AI) · July 17, 2026When Does Muon Help Agentic Reinforcement Learning?