3 gaps in states’ AI governance for public benefits
AI adoption is outpacing governance and evaluation. We should build a feedback loop.
At Nava Labs we’ve been thinking a lot about AI adoption and policy in the public benefits space: where states stand and what “good” might look like. Our running hypothesis is that “good” should be a loop that connects governance, adoption, and evaluation.
But before we get to good, let’s talk about the current state.
The current state
There’s a lot of conversation about “AI readiness,” including what leadership, governance, and infrastructure should be in place for public-sector agencies to deploy AI tools in benefit programs like SNAP. At this stage, both federal and state legislative guidance is thin (only three states have AI laws that explicitly impact AI use in public benefit programs), and will likely stay that way. Instead, we’ve seen many states turn to governors’ executive orders, AI task force reports, and administrative policies to add adequate controls and guardrails around how AI can be deployed, seeking to prevent harm.
Once the policies and infrastructure are in place, public benefit agencies often decide they’re ready for their first AI pilots, often starting with internal-facing tools that aim to reduce caseworker burden. Many states have begun piloting numerous AI tools including policy chatbots, case note generators, document quality verifiers, and the like — tools Nava has had a hand in delivering.
And if the state agency has the right infrastructure and capacity, or if they’re working with a trusted vendor, hopefully they’re also doing some version of the third step: AI evaluation.
With all three steps in place, the process would look something like this:
This is a good start, and agencies who’ve developed some form of this process should applaud their progress — then push a little further.
And while most states establish governance before adoption, it doesn’t always happen in a linear fashion. States who fail to govern, or who establish strict guardrails that don’t allow for safe experimentation, may see shadow AI adoption by frustrated employees looking to reduce workload burden behind the scenes.
In states who’ve established some form of generic AI governance as the first step, we’re also starting to see a pattern where AI adoption is increasing faster than the two things that should keep it accountable: evaluation and governance. AI technology is evolving faster than the other two can keep up.
Current gaps, and how states might address them
A few common problems we’ve identified across many states with this current process:
1. Evaluating technical performance but not impact
Many states are missing a critical step in the above process: AI impact evaluation. Any AI tool deployed for any reason is typically evaluated by the deployer for a certain period of time to ensure quality, safety, reliability, and other performance metrics.
But human service agencies don’t care solely about the performance of the tool — they care whether it’s making a difference in the program performance metrics they’re expected to meet: Did it reduce the SNAP Payment Error Rate? Did it decrease churn among participants subject to work requirements? Did it reduce denials for failure to return verifications? These are the questions state agency leaders are responsible for answering to their legislatures when it comes time to fund their programs.
Despite this need, we’ve found that states are often lacking in AI impact evaluations, which are the types of evaluations that actually answer those questions. This may be due to lack of data infrastructure, lack of appropriate staff and resources, or lack of understanding around how to conduct impact evaluations. Whatever the reason, this is an area we’ve seen for further growth in many states.
We believe evaluations — including impact evaluations — are a key component of AI adoption in the public sector that should be prioritized. These evaluations help ensure the tool not only works but it actually meets the outcomes set by the agency’s original goals. We believe at a minimum, all AI tools deployed in public assistance programs should be able to answer these three questions:
Does it reduce caseworker burden?
Does it improve client access?
Does it cause any unintended harm?
For states who are struggling to get started with impact evaluations, Nava has put together some great resources you can start with.
2. Not sharing learnings across deployments
Another common problem we’ve seen is that the different AI tools are being piloted by different vendors, across different programs, and by different agencies. What one agency learns in their user research and evaluations rarely makes it to the next, causing duplicated work and inconsistent deployments that don’t build on each other.
We believe states that want to build AI tools effectively, quickly, and consistently in their benefit programs should have a system in place to share learnings across different deployments. This ensures that the same mistakes don’t get repeated from program to program.
There are now approximately fourteen states that are publishing or are required to publish inventories for either “high risk” or all AI use cases within government agencies. We think an AI inventory — at a minimum for “high risk” use cases — is a good start, but we think this can go even further.
First, states need to ensure that the AI use case inventory process isn’t so burdensome and restrictive that their agencies are afraid to populate it. This is a pattern we’ve seen repeated a few times across states: inventories that require detailed submission, review, and approval. They ultimately scare agencies away from disclosing certain use cases they think will take too long to approve, or won’t be approved at all. We would caution that a burdensome inventory process can be equivalent to no inventory at all (and uses significantly more resources to set up).
Second, we believe that good AI inventories don’t just include a list of the use cases, but also share their evaluation results and any lessons learned. If other agencies are using the inventory as a way to understand what exists and how they can deploy it, these results will be critical for helping them make informed decisions and ask potential vendors the right questions.
A good AI inventory might also include space for the following information:
Agency owner of this deployment, or who to contact for more information
User research and evaluation findings or how to access them
Best practices from the deployment. For example:
“We found a good way to ask users about consent to using AI. Here’s the pattern and language we used…”
“We were able to safely engage users of this particular demographic by doing…”
“We developed a way of ensuring meaningful human review of the final output. Here’s what we did…”
Negative outcomes or harms discovered that should be shared with others developing a similar tool
3. Missing mechanisms for policy revision
We’ve found that in states that have written AI policies, almost none have a process in place to review and revise that policy. Once published, there’s often no clear owner reviewing the AI policy to make sure it did what it was intended to do. Did the guardrails protect who they were meant to? Are there any additional ones needed? AI technology is changing rapidly. What once made sense at the time of publishing may no longer be sufficient to protect against the capabilities of newer models or agentic AI.
We believe that AI governance should be revisited regularly, especially early in a state’s experimentation with AI. While legislation is a lot harder to change, there are more flexible policy vehicles like statewide executive orders and government-wide policies. This is the level at which states can maneuver more quickly and effectively. Good AI governance will only get better over time, and this can be enhanced and expedited if it operates as a loop — where governance feeds adoption, which feeds evaluation, which feeds back to governance and adoption.
We believe forming a strong loop from the start is what will separate states that simply adopt AI from states that get better at it.
Next steps
If your state is just starting out on your journey to AI adoption, or if you have a process and you’re looking to refine it, consider how you’ll build a loop. This could look like many things, including:
Build up your agency’s data infrastructure to allow for better impact evaluations
Write evaluation criteria into your RFPs
Create a post-deployment review of every AI use case that helps identify best practices and negative outcomes that might warrant adjustments to policy
Develop an AI inventory that also includes evaluation and learning data that’s shareable at least across your agencies, if not publicly
Codify and share any best practices that have proven effective in your program
Set a regular review cadence of your AI policies
Designate one clear owner of an AI policy whose job is to revise based on learnings
Consider joining feedback channels and working groups with other states that are interested in cross-state sharing of learnings
If you’re also thinking about this topic, we’d love to hear from you about your experience. Is there anything we’re missing? Let us know in the comments, or email us at labs@navapbc.com.





