Binglin Zhou, Peng Shi, Ryo Kamoi +2 more
Checking a scientific claim against the paper it came from usually means reading a figure. The number is in a chart, the comparison is in a table, and the qualification that matters is in a caption. Models attempting this fail in three distinguishable ways: they cannot find the decisive visual evidence, they misread structured visuals when they do find them, and they fail to combine what they saw with what they read.
Those are separate failures needing separate remedies, which is the argument for a tool-augmented approach. Locating a panel, extracting values from a chart and reasoning over the result are different operations, and a single forward pass asks one mechanism to do all three at once.
The stakes differ from ordinary fact-checking too. A claim that is right except for a factor of ten reads as plausible, and catching that requires actually reading the axis rather than recognising the general shape of the figure.
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our…
RoboTTT: Context Scaling for Robot Policies
arXiv (cs.LG) · July 16, 2026MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
arXiv (cs.AI) · July 16, 2026SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
arXiv (cs.LG) · July 16, 2026Online Neural Space Time Memory for Dynamic Novel View Synthesis