AI for public authorities: use cases, requirements and choosing the right model
How artificial intelligence is already being used in public authorities
Language models are already supporting public-sector work in a number of areas:
Experience with widely available commercial LLMs such as Anthropic’s Claude, OpenAI’s GPT, Google’s Gemini, Mistral AI’s Mistral and Aleph Alpha’s Luminous has understandably increased interest in adopting AI tools more widely across the public sector. There is clear scope for integrating them more deeply into administrative workflows.
What challenges does AI present for public authorities?
This leaves public-sector decision-makers facing a dilemma: language models are advancing faster than their suitability for administrative use can be assessed with any degree of certainty. In Germany, individual public authorities also have considerable freedom to decide which AI tools to use, how to use them and on what scale. Without binding nationwide rules and standards, this risks creating a fragmented landscape of different technologies, practices and levels of oversight.
It also raises a fundamental question: which language model is suitable for which task? There is no single model that can perform every administrative task equally well. The requirements involved in answering a straightforward enquiry from a member of the public are very different from those associated with complex documents or specialist administrative procedures. The relevant question is therefore not which model is regarded as the best overall. What matters is which model best fits the authority’s particular requirements and intended use.
One way forward is to compare language models against the tasks they will actually be expected to perform.
What a meaningful model comparison needs to assess
This is particularly important because conventional model comparisons have significant limitations. Most benchmarks focus on capabilities such as understanding English-language texts, solving mathematical problems or demonstrating general knowledge. They reveal far less about the issues that matter in day-to-day public administration:
- When responding to an enquiry, does the model invent information that has no basis in legislation or regulations (hallucinating)?
- Can it work reliably in German, including with administrative terminology and complex procedural requirements?
- Does the provider explain clearly which data were used to train the model?
- How is information generated or processed by the system protected?
Established language models such as Claude, GPT and Gemini tend to perform well in general benchmarks. That does not automatically make them suitable for use by public authorities. A well-known name and strong overall performance are no substitute for an assessment based on the realities and requirements of public administration.
Public authorities therefore need evaluation criteria of their own – in other words, benchmarks designed specifically for LLMs and AI tools used in the public sector.
AI for public authorities: performance and governance as benchmarks
Any comparison of models designed for public administration must examine two separate but equally important dimensions: performance and governance.
1. Performance: how well does the model do the job?
Performance is concerned with the model’s practical capabilities: how well it completes specific administrative tasks. Relevant criteria include:
- Summarising: producing accurate, well-structured summaries of decisions, specialist documents and court judgments.
- Answering questions: providing precise answers based on the documents and information made available to the model.
- Extracting topics: searching documents, identifying relevant subject matter, categorising content and assigning suitable keywords.
These tasks depend on several underlying capabilities:
- Command of administrative language: understanding the terminology and conventions used by public authorities.
- Factual accuracy: producing correct results without introducing fabricated information.
- Control of style and register: communicating consistently and in a tone appropriate for official correspondence.
2. Governance: can the model be deployed responsibly?
Governance is concerned with the conditions under which an LLM can be used responsibly in the public sector. It covers the traceability of outputs, data protection, information security and compliance with regulatory requirements. The following criteria are therefore central to the transparent, compliant and legally robust use of AI tools:
How often the model introduces claims that are not supported by the source material or misrepresents the information provided.
Whether its outputs are compatible with democratic principles and fundamental values.
How efficiently the model uses computing resources.
How openly the provider communicates information about the model’s architecture, its training data and the terms governing its use.
Governance also encompasses the following areas:
A useful comparison must assess performance and governance separately while giving both equal weight. After all, even the best model is of little use if the conditions necessary for it to be used safely and effectively are not in place.
Questions of security and data location also relate directly to a broader principle for the public sector: digital sovereignty.
Data sovereignty – the key to digital sovereignty
Digital sovereignty is the ability of the state and public administration to shape, control and use digital systems, processes and technologies on their own terms. Achieving it requires reducing critical technological dependencies on suppliers outside Europe. This is not about isolating European public administrations from the wider technology market. It is about actively developing and adopting trusted European alternatives and open standards. Doing so helps preserve the state’s ability to act independently and strengthens democratic resilience in the digital age.
Read more here about digital sovereignty.
The model is only part of the picture
Assessing an LLM for use in public administration involves more than comparing the capabilities of the model itself. The technical, organisational and security arrangements surrounding its deployment are just as important.
These include the chosen deployment model, the available IT infrastructure and interfaces, the authority’s own expertise and knowledge, the budget available for operation and further development, and internal requirements relating to data protection and information security. Operational resilience also matters. AI-based services must continue to function reliably even when large numbers of users access them at the same time.
Understanding the available deployment models
Public authorities broadly have two deployment options:
A hybrid arrangement combining both approaches is also possible. Lower-risk tasks can be handled in the cloud, while particularly sensitive information is processed exclusively on premises.
Protection requirements The appropriate deployment model depends on the level of protection required for the data involved. The BSI’s protection requirement categories – normal, high and very high – determine the requirements that apply to infrastructure, encryption and access controls. General enquiries from the public may, in some circumstances, be suitable for processing in a certified government cloud. Information with a high protection requirement will generally need to be processed within the authority’s own secure infrastructure.
Regardless of the deployment model, retrieval-augmented generation (RAG) can be used to incorporate an authority’s own knowledge into the system. When a user submits a query, the system first searches an approved collection of documents – such as legislation or internal guidance – and uses the relevant passages to inform its answer. This can reduce hallucinations, make outputs easier to trace back to their sources and be implemented both on premises and in the cloud.
Bundesdruckerei: developing practical AI for public authorities
Since 2023, Bundesdruckerei GmbH has been supporting Germany’s federal administration through its AI Competence Centre (KI-KC). Its work helps federal authorities use AI on their own terms and develop applications grounded in real administrative needs. Commissioned by the Federal Ministry of Finance, the centre works with a number of federal ministries on projects examining how AI can be used in day-to-day public administration.
The emphasis is on solutions designed specifically for public authorities. Current areas of focus include language models, computer vision and anomaly detection, all aimed at making administrative processes more efficient and effective. One of these projects is MÖVE – Evaluating models for public administration.
Bundesdruckerei’s MÖVE project: an LLM benchmark for public administration
MÖVE is an assessment framework that brings technical performance and governance requirements together within a single system. It is designed around the practical realities of public administration.
- The project aims to develop a benchmark based on tasks and datasets that reflect the actual work of public authorities.
- It evaluates performance-related capabilities such as summarising, extracting topics and answering questions, as well as governance factors including hallucinations, transparency, sustainability, politics and values.
- MÖVE enables authorities to compare models systematically, assess their suitability for internal documents and workflows, and weigh up the associated opportunities and risks.
The result is an LLM benchmark that offers practical guidance on selecting appropriate AI models for public administration – securely and transparently.
Further projects focusing on public-sector AI and data sovereignty
Questions and answers about AI for public authorities (FAQ)
The decision should be based on the task the system will perform – such as communicating with the public or supporting internal work – as well as the available IT infrastructure and the authority’s particular governance and security requirements. A benchmark designed for public administration, such as the MÖVE framework, makes it possible to compare language models and the AI tools built around them systematically.
This requires a structured evaluation against clearly defined criteria covering technical quality, IT operations and governance. Documented test results, standardised benchmark reports and audit evidence provide a transparent and auditable basis for the choice of model. They also help demonstrate that the system is being used in accordance with the relevant legal requirements.
Generated text and analysis can be evaluated against reference material. Assessments will typically combine automated checks with expert human review. Relevant tests include plausibility checks and reviews of factual accuracy, completeness and consistency. Regular updates help ensure that performance remains stable over time.