Laurens Samson, Iva Gornishka, Gossa Lô +2 more
Governments are putting language models into public administration, where the requirements are not the ones benchmarks usually measure. This work builds an evaluation suite for Dutch governmental use, developed with domain experts from a major Dutch municipal organisation rather than assembled from existing leaderboards.
The dimensions came out of an advisory board process, user research and a survey of people actually using a civil-servant chatbot. Six emerged: factuality, honesty, social bias, energy consumption, cost, and training-data transparency — notably placing environmental and financial cost alongside quality, as things a public body has to answer for. Those were operationalised into a benchmark covering more than 30 multilingual and Dutch-specific models.
The headline finding is that no model wins on everything and the trade-offs are structural: higher quality consistently costs more in both environmental impact and money. Bias, they report, moves largely independently of both — so it cannot be bought down by spending more.
Scope worth keeping in view: this is Dutch-language governmental use, with dimensions chosen by one municipal organisation. That specificity is the point of the work, and it is also the reason the abstract does not claim the same six dimensions or the same trade-off curve would fall out for another language or another arm of the state.
Written by the Hevolve AI agent from this paper's abstract, and reviewed by a person before publication. The abstract is quoted below so you can check it against the source.
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey…
An Exam for Active Observers
arXiv (cs.LG) · July 17, 2026PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment
arXiv (cs.CV) · July 17, 2026CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception
arXiv (cs.CV) · July 17, 2026VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.