Yasheng Sun, Zezi Zeng, Yifan Yang +4 more
Anyone who has revised a paper knows the tedium of fixing figures: relabeling parts, shuffling panels, restyling for a reviewer. SciDiagramEdit goes after automating that from a plain-language instruction. The hard part is that a scientific figure isn't a photo, it's a dense infographic where arrows, plots, captions, and schematics follow a tight visual grammar to make an argument, so edits must respect that structure.
Their approach has two pieces. One is a benchmark built by mining before-and-after figure pairs from arXiv version histories, so the edits reflect what authors actually changed. The other is an agent that works on the figure's editable vector source (so you can inspect and tweak individual pieces alongside it) and improves through skill evolution, where a proposer keeps refining the agent's skill description from its own traces. Over several rounds, edit accuracy on held-out figures rises.
This comes only from the abstract, which reports the trend rather than hard numbers, so the paper will have the concrete results.
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight…
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
arXiv (cs.CV) · July 17, 2026Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
arXiv (cs.LG) · July 17, 2026PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
arXiv (cs.LG) · July 17, 2026A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing