Pengcheng Zhou, Xuanyu Liu, Yanchen Yin +4 more
Ask a vision-language model where a photo was taken and it does well when a famous landmark is in frame. The trouble is it leans too hard on those landmarks, ignoring quieter clues like road markings, plants, or building styles, and sometimes it invents connections that aren't there. HoloGeo is about weaning models off that crutch.
The authors first build tools to measure the problem: two metrics for how strongly and how harmfully landmarks sway a model, plus a benchmark called LandmarkBias-3K. Then they train a model to reason from spread-out evidence rather than one obvious cue, using a dataset annotated with step-by-step, bias-free reasoning chains and rewards that push it to weigh several visual cues at once. They report holding steady on standard geo-localization tests while doing notably better on their bias-focused benchmark.
Since this comes from the abstract, the paper carries the metric definitions and the full comparison against other open models.
Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization. To systematically investigate this issue, we first design two quantitative metrics, Bias Intensity (BI) and Bias Harmfulness (BH), to characterize the impact of landmarks exerted on model reasoning,…
Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting
arXiv (cs.CV) · July 17, 2026Adaptive Contrast Enhancement and Optimised Feature Matching for RootSIFT-Based Palm-Vein Recognition
arXiv (cs.AI) · July 17, 2026HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection
arXiv (cs.AI) · July 17, 2026JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models