Sushant Gautam, Vajira Thambawita, Michael A. Riegler +2 more
In medical AI that reads both images and text, getting the answer right isn't the whole job; a clinician also has to trust the reasoning behind it. This paper steps back from the leaderboard and looks at nine systems from a gastrointestinal endoscopy challenge to see which design choices actually produce reliable behavior, not just high scores.
The honest finding is that a good answer and good reasoning don't always travel together. Lightweight fine-tuning of pretrained models scores well on the challenge, but those wins don't reliably turn into faithful, complete clinical explanations. Systems that force structured reasoning and point explicitly at their evidence behave more dependably across question types, though the authors stress this is a correlation they observed, not something isolated with controlled ablations. They argue for evaluation that moves past word-overlap scores toward evidence-linked explanations, leakage-aware data handling, and basic robustness and calibration checks.
This is a retrospective analysis summarized from the abstract, so the paper carries the per-system detail and caveats.
Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of pretrained backbones provides strong challenge performance, but answer-level gains do not consistently translate into faithful and complete clinical…
Beyond Unfolding: 60x Faster One-Stage Unmixing for Closely-Spaced Infrared Small Targets
arXiv (cs.AI) · July 17, 2026Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids
arXiv (cs.CL) · July 17, 2026BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC
arXiv (cs.LG) · July 17, 2026CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors