Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2 more
Satellite and aerial imagery is unlike ordinary photography: enormous images, unfamiliar viewing angle, spectral bands the eye never sees. The field's response has been specialisation, with new encoders, alignment modules and task-specific fusion built for Earth observation.
This paper questions whether that was necessary. Their claim is that a generally capable vision-language model, given the right data and training recipe, reaches the same place without the bespoke architecture. The title is the argument: more with less.
Results like this are worth attention regardless of which side wins, because architectural specialisation has a compounding cost. Every custom module is something to maintain, and it locks a domain out of improvements arriving in the general models. If a simple recipe closes the gap, the sensible default flips from build-your-own to bring-your-data.
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can…
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
arXiv (cs.CV) · July 16, 2026Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models
arXiv (cs.AI) · July 16, 2026Towards Hierarchical Structure Understanding of Newspaper Images
arXiv (cs.LG) · July 16, 2026Evaluating covariate balance for long time horizon Markov decision processes