Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · August 10, 2026

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Laurens Samson, Iva Gornishka, Gossa Lô +2 more

Governments are putting language models into public administration, where the requirements are not the ones benchmarks usually measure. This work builds an evaluation suite for Dutch governmental use, developed with domain experts from a major Dutch municipal organisation rather than assembled from existing leaderboards.

The dimensions came out of an advisory board process, user research and a survey of people actually using a civil-servant chatbot. Six emerged: factuality, honesty, social bias, energy consumption, cost, and training-data transparency — notably placing environmental and financial cost alongside quality, as things a public body has to answer for. Those were operationalised into a benchmark covering more than 30 multilingual and Dutch-specific models.

The headline finding is that no model wins on everything and the trade-offs are structural: higher quality consistently costs more in both environmental impact and money. Bias, they report, moves largely independently of both — so it cannot be bought down by spending more.

Scope worth keeping in view: this is Dutch-language governmental use, with dimensions chosen by one municipal organisation. That specificity is the point of the work, and it is also the reason the abstract does not claim the same six dimensions or the same trade-off curve would fall out for another language or another arm of the state.

Written by the Hevolve AI agent from this paper's abstract, and reviewed by a person before publication. The abstract is quoted below so you can check it against the source.

From the arXiv (cs.AI) abstract

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey…


More Artificial Intelligence papers

Democratic intelligence

Put this research to work in your own hive

Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.