The Missing Ingredient in Government AI: Proof That It Works

By Jeremie Ponak

For the past few years, government modernization advocates like Jennifer Pahlka have made the case that rebuilding civic trust requires government to actually deliver results. If public services work — benefits arrive on time, forms aren't nightmares, the DMV doesn't ruin your afternoon — people trust the institutions behind them. AI could accelerate that: faster benefits responses, better accessibility, smarter infrastructure. It could also undermine it with incorrect determinations, opaque denials, and confident misinformation delivered at scale.   

During my time at the Rockefeller Foundation and Columbia University, I’ve spoken with many government and civic technologists. Despite preconceived notions of public servants as cautious and resistant to change, most are embracing AI. In fact, there's been tremendous growth in AI applications in the public sector: the federal government tracked over 2,100 AI use cases in 2025 alone, tripling from the year before and likely to grow significantly by the time of this writing. But agencies are still hesitant to deploy LLMs in the one place where they could have the most impact: front-line, constituent services delivery. Every state agency I’ve spoken to is actively exploring, testing, or piloting AI in order to make them more efficient, adaptable, and effective, but a lot of that is focused on drafting tools, summarization, and employee productivity.

We are facing historic lows (red arrow) for trust in government, 19% in 2025.Source: Pew Research Center.

Early constituent-facing deployments have surfaced the challenge. Without robust pre-deployment evaluation, even well-intentioned tools can provide incorrect information. And the trust gap is significant: 50% of residents say they're uncomfortable with government agencies using AI for public services, up from 45% the year before. So while AI usage is growing, confidence isn't. Part of the reason is straightforward: there's limited credible ways to measure whether these systems actually work for the tasks the government needs them to do.

Could we close the trust gap by putting leading AI models to the test against real policy and eligibility questions, and making their performances available publicly, evaluating them based on their accuracy, completeness, or reliability? That’s what I’ve been testing, with open AI benchmarks on the public sector’s most important problems: benefits eligibility, service delivery, healthcare. I started experimenting with SNAP eligibility, one of the most rule-intensive and high-stakes domains in government service delivery. Over 42 million Americans rely on SNAP. The policy is complex, it just went through a major legislative overhaul, and it's exactly the kind of thing agencies are using AI for. This is the work that matters most to constituent experience, and it's also the place today where a wrong answer directly harms someone.

 

How can AI benchmarks help?

Benchmarking AI isn't a novel concept. In the research community, benchmarks like GLUE, SQuAD, and MMLU have been foundational, giving developers clear targets, creating comparability across models, and measurably driving progress. Today, benchmarks are being used to test LLM’s abilities to perform jobs, from finance to taxes to legal advice. The improvement cycle is essentially: release benchmark → models compete → performance improves → release harder benchmark. It works. Leading models can code, do math, and design beautiful slide decks, and are getting better day by day.

Most benchmarks, however, test general capabilities like reading comprehension, reasoning, code generation. No public-sector-specific AI benchmark exists, and that means every agency deploying AI is running its own evaluations with no way to compare notes. None of them test whether a model can calculate SNAP income eligibility for a household of four with a dependent care deduction and a state-specific Broad-Based Category Eligibility (BBCE) waiver. None of them test whether a model correctly handles the difference between general work requirements (exempt at age 60) and ABAWD time limits (applies through age 64 under new law). None of them test whether a model refuses to help someone commit benefits fraud.

That’s not to say that governments haven’t been evaluating their AI deployments; on the contrary, public sector deployment tends to be careful and consistent, partnering with experts from Code For America to Propel to understand how LLMs perform. But there's no shared infrastructure for comparing results across states or agencies, no common format, no standard rubric, no public repository. Evaluations, where they exist, are tied to a specific moment and a specific deployment, with no mechanism to build on what others have learned. Why?

 

Which questions might public sector benchmarks help us answer?

  • Which models handle policy complexity best? Not all LLMs are equal on rule-intensive tasks. 

  • Where do models fail in ways that matter? General accuracy numbers hide the important patterns. Knowing where models fail is as useful than knowing how often they fail.

  • How quickly do models absorb policy changes? Policy changes constantly, and any evaluation framework needs to track whether models keep up.

  • Do models behave safely on adversarial inputs? People can ask sensitive and wrong-headed questions all the time – and agencies need to know which models are better at steering conversations towards productive outcomes.

  • Is there a meaningful accuracy/cost tradeoff? As agencies scale AI usage, compute costs matter. Determining which model does a given task most efficiently, as well as accurately, will be an evaluation criteria.

 

What we found when we started testing

We built a pilot benchmark: 30 expert-validated questions across income eligibility, work requirements, immigration status, household composition, benefit calculations, and more. Each question has a ground-truth answer sourced line-by-line from USDA FNS, the Code of Federal Regulations, and state agency guidance. We then ran a set of leading models against these questions using a rubric across five dimensions: factual accuracy, policy adherence, completeness, safety, and tone.

The results were instructive, even on a small set of questions, testing smaller models (GPT 4o mini, claude 3 haiku, and gemini 2.5 flash). On straightforward questions — basic income thresholds, standard deductions, simple household scenarios — models generally performed well. This tracks with what Dave Guarino, formerly of Propel has been finding in his work testing LLMs on SNAP policy: for the basics, they've gotten surprisingly capable. (Dave's team has done particularly useful work using AI to help SNAP recipients diagnose and restore lost benefits.)

But the cracks showed up fast when complexity increased. Models struggled with recent policy changes, particularly the One Big Beautiful Bill Act's (OBBBA) rewrite of work requirements and immigration eligibility. They struggled to distinguish federal rules from state-specific variants. And they struggled with the layered edge cases that real applicants actually face: mixed-status families, elderly individuals navigating the distinction between general work requirements and ABAWD, refugees who've adjusted to LPR status. 

Using GPT-4.1 Mini without policy-specific RAG (meaning the model answered from training data alone, as a constituent would experience asking ChatGPT) accuracy averaged around 40%. This is an early baseline with a smaller model, not a representation of production deployment. But it matters because this is how millions of people are already using these tools.

Some models confidently provided wrong answers — the kind that would lead a caseworker or an applicant to make a bad decision. We plan to publish detailed model-by-model results and expand to a broader set of models in the coming weeks.

 

The policy moment

In some ways, the massive uptick of AI use in government is perfectly timed for this moment of policy complexity. HR1 rewrote significant portions of SNAP policy in mid-2025, leaving states with a tight window to adapt and increased reason to tackle downstream issues like payment error rates. And parsing policy is a bipartisan issue, with as many conservatives as progressives looking for ways to simplify decades of accumulated law, whether that's making SNAP delivery more precise, speeding up environmental reviews, or making cybersecurity best practices portable across agencies.

Timeline of the HR 1 (OBBBA) changes to SNAP, medicare, and medicaid. Source: Washington State Department of Social and Health Services

LLMs can help with this — and all over the country, they already are. Tools like RegLab’s STARA are helping governments navigate legislative complexity for permitting reform, and organizations like OpenPolicy (where I work) are building AI pipelines to help companies and agencies parse regulatory change in real time. But as long as each effort tackles one problem in isolation, with its own evaluation and no shared evidence base, the government will never build the capacity to respond quickly enough. That's where shared benchmarking infrastructure comes in — the same methodology that evaluates SNAP eligibility can extend to Medicaid, to housing, to any domain where policy precision matters.

 

Who this is actually for

There's a tendency to frame public-sector AI benchmarks as primarily an accountability mechanism – which they are, but they also have practical value for a whole host of actors. Agencies can stop reinventing evaluation from scratch, starting with a common reference point for procurement and risk review. Vendors and model providers understand what "good enough for government" actually means, and have a benchmark to optimize against. And finally, developers and civic technologists get open datasets that they can use to test their own applications before deployment. 

 

What's next

This is early-stage work, but we're actively building this out with All Tech Is Human's network and in conversation with government partners. If you're interested in the work — as a collaborator, funder, or domain expert — we'd like to hear from you.

Two specific asks:

Are you working in government and deploying or evaluating AI, especially in food policy, benefits administration, or healthcare? We want to work with experts in their respective domains to test the most common and also the hardest policy and eligibility questions. Your domain knowledge is what makes this credible.

Do you have experience or interest in building benchmarks or AI evaluation for the public sector? Whether you're a researcher, a civic technologist, or someone who's been thinking about this from another angle — we want to connect.

Reach out to jeremie@civbench.org or fill out this form. We'll be publishing expanded results and full methodology soon. All Tech Is Human has been exploring public interest technology projects, such as this collaboration with Jeremie Ponak.

Previous
Previous

The Agentic Shift: Navigating the New Human-AI Workflow

Next
Next

Jen Weedon to Join All Tech Is Human’s Braintrust