Turning public values into benchmarks
The “Grip on LLMs” project built a public-sector evaluation framework with Dutch municipal experts, an advisory board and research involving users of a civil-servant chatbot. It measures six dimensions: factuality, honesty about uncertainty, social bias, energy use, cost and training-data transparency. More than 30 multilingual and Dutch-specific models were compared.
The authors report that no single model led every dimension. Higher quality generally carried greater environmental impact and financial cost, while bias was largely independent of both. Factuality and honesty also behaved as distinct properties: a model that answers more questions correctly does not automatically admit uncertainty more reliably.
Procurement implication
Governments should define the failure cost for each workflow—citizen answers, internal search or translation—before selecting a leaderboard winner. Consequential answers need evidence and abstention tests. Energy and price should be compared at equal workload and service levels, with language-specific bias checks and appeal paths.
Research boundary
This is a preprint centered on Dutch and collaboration with one major municipal organization. Model versions, prices, electricity mix and hosting can change results, and the framework cannot be transferred unchanged to another country's law and language. Its useful contribution is making technical and public-value trade-offs visible in one selection process.