Xiao Lin, Xiaohu Huang, Kai Han
Multimodal language models are getting better at spatial reasoning, and a common boost is to feed them priors from a pretrained vision model, say a depth or segmentation expert. The usual setup picks one such expert. The observation driving this paper is that different experts are good at different things, so committing to a single one leaves value on the table.
ViPS gathers several. An Efficient Prior Proxy produces multiple foundation-model priors without much added inference cost, and a Dynamic Prior Fusion module decides how to blend and inject them depending on the task at hand rather than mixing them uniformly. The authors report state-of-the-art results across a range of spatial reasoning and 3D understanding benchmarks.
The abstract stays at the level of components and headline claims, so the paper and project page are where you would check how much each expert actually contributes.
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a…
CRISP: Constrained Refinement via Iterative Squeezing Process for Robust Medical Image Segmentation under Domain Shift
arXiv (cs.LG) · July 16, 2026Data Driven Block Replacement Scheduling
arXiv (cs.CV) · July 16, 2026Divergent Gaze Patterns in Artistic Viewing: Spatial and Temporal Signatures of Attention Across Autistic Individuals, Artists, and Neurotypical Observers
arXiv (cs.CV) · July 16, 2026Structural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-Identification