Jiarui Zhang, Muzi Tao, Shangshang Wang +3 more
Human seeing is not a snapshot. Your eyes move constantly, and where they go next depends on what you have provisionally concluded from what you have already seen. Psychophysics has argued for decades that this closed loop is not a detail of vision but essential to it.
Multimodal language models are handed a finished image and asked about it. Whether they do anything resembling active observation is an empirical question, and the authors point out that existing vision-language benchmarks cannot answer it, because a benchmark that presents one static image has already removed the behaviour it would be testing.
ActiveVision is built to make the loop measurable. The design lesson generalises beyond vision: if your evaluation removes the conditions under which a capability would be exercised, a high score tells you nothing about whether the capability is there.
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active…
Revisiting data-driven dynamic security assessment with a tabular foundation model
arXiv (cs.AI) · July 17, 2026Rethinking Quantum Continual Learning with Quantum Fisher Information
arXiv (cs.LG) · July 17, 2026CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach
arXiv (cs.LG) · July 17, 2026Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints