Government (Public Sector)

Government agencies operate under a set of constraints that make agentic AI deployment both more challenging and more consequential than in commercial environments. Mission-critical services cannot fail. Decisions affecting citizens require explainability and legal defensibility. Legacy infrastructure, some of it decades old, must integrate with modern AI systems. And public trust — once lost — is extraordinarily difficult to rebuild.

Against this backdrop, 2025 marked a turning point: agentic AI moved from pilot projects in government to authorized, production-scale deployments. The implications are significant for any public sector organization still developing its strategy.

What Deployments in 2025 Demonstrated

The State Department’s CIO office deployed agentic AI across legacy systems that had resisted previous modernization efforts, using agents as an integration layer that could query across disparate databases without requiring the underlying systems to be replaced. This model — agents as connectors rather than replacements — has proven to be the practical path forward for agencies whose infrastructure cannot be rebuilt on an accelerated timeline.

USPS deployed agents to deliver personalized customer experiences at scale, handling service inquiries, package tracking escalations, and delivery preference updates through conversational agents that maintain context across interactions. For an organization serving hundreds of millions of customers, this represents a structural shift in what “personalization” means at public-sector scale.

The CIA’s enterprise automation initiative demonstrated that agentic AI can operate effectively across diverse, siloed databases — synthesizing information across systems that were never designed to communicate with each other. The security architecture required to enable this while maintaining need-to-know access controls became a template for intelligence community deployments.

Oklahoma’s state government provided one of the most operationally significant examples: AI agents triaging thousands of Security Operations Center alerts daily, reducing analyst fatigue and accelerating response times to genuine threats. In cybersecurity contexts, the volume of alerts has long exceeded human capacity to review them; agents that can classify, prioritize, and route alerts represent a meaningful capability upgrade.

The Regulatory Framework Is Clarifying

Two April 2025 OMB memos formally empowered Chief AI Officers as change agents within federal agencies, giving them authority to direct AI strategy, set standards, and hold programs accountable. This structural change matters: it means agentic AI programs now have designated executive sponsors with defined authority.

OMB M-25-22 (September 2025) established a critical boundary: vendors are prohibited from using non-public government data to train models. This is not simply a data governance policy — it is a procurement and architecture requirement. Any agentic AI platform deployed in a government context must demonstrate, at the infrastructure level, that operational data does not flow back to model training pipelines. Contracts and system designs must reflect this constraint explicitly.

ServiceNow AI Agents achieved FedRAMP High Authorization at Impact Level 5 — the most rigorous federal cloud security standard. This certification signals to the broader government market that enterprise-grade agentic platforms can meet the security requirements of even the most sensitive civilian and national security workloads. Organizations evaluating platforms should treat FedRAMP authorization as a baseline requirement, not a differentiating feature.

GAO’s analysis projects that AI agents will be embedded in one-third of enterprise software applications by 2028. For government agencies, this means agentic AI is not an optional modernization initiative — it is the direction the enterprise software ecosystem is moving, and agencies that do not build the governance capacity to manage it will find agents appearing in their environments through routine software procurement, without the oversight structures to govern them effectively.

Design Requirements for Government Contexts

Reversible resilience. Agents operating in government workflows must be designed with rapid rollback capability from the outset. When an agent error propagates through a high-volume workflow — incorrect eligibility determinations, misdirected alerts, erroneous correspondence — the ability to identify the error, halt the agent, and reverse its actions quickly is not a nice-to-have. It is a mission requirement. Rollback procedures should be tested before deployment, not developed in response to an incident.

Explainability for public accountability. Government decisions are subject to FOIA requests, congressional inquiries, and judicial review. Any agentic system contributing to a material decision must be able to produce a plain-language explanation of its reasoning — one that a citizen, a journalist, or a congressional staffer can understand. This requires agents to log not just their outputs but their intermediate reasoning steps and the evidence they drew upon.

Human sign-off on citizen-affecting decisions. Across all 2025 government deployments, one principle held constant: decisions that materially affect individual citizens require human authorization before execution. Agents prepare, analyze, and recommend; credentialed officials decide. This division of labor is not simply a policy preference — it reflects the legal accountability framework within which government agencies operate. Agents that usurp final decision authority create liability that no technology vendor can indemnify.

Access control at the data level. Government agents routinely need to query across systems that contain information classified at different sensitivity levels. Access controls cannot be managed through agent instructions alone — they must be enforced at the data layer, with the agent’s tool-calling permissions audited continuously. An agent that can be prompted to access data beyond its authorized scope through adversarial inputs represents a security vulnerability, not just a policy violation.

The Cautious, Incremental Path Forward

The pattern across successful 2025 government deployments is consistent: start with internal-facing applications — employee copilots for document drafting, information retrieval, report synthesis — before deploying agents in citizen-facing or decision-making roles. This sequence allows agencies to build operational familiarity, identify failure modes, and develop oversight practices before the stakes are highest.

Agencies that have tried to accelerate past this sequence — deploying autonomous agents in citizen-facing roles before establishing oversight infrastructure — have encountered the predictable consequences: public trust incidents, congressional scrutiny, and forced rollbacks that set programs back further than a deliberate pace would have.

The technology is ready for government use. The question for each agency is whether its governance infrastructure is ready to use the technology responsibly.

A Four-Stage Progression

Stage 1: Internal productivity applications (months 1–12)

Deploy agent copilots for employees: document drafting, policy research, meeting summarization, information retrieval across internal systems. These applications have no direct citizen impact, allow staff to build operational familiarity with agent behavior, and produce the incident reports and failure taxonomy that will inform governance policy for citizen-facing applications. Success criteria: 80%+ user adoption, zero uncontained incidents, documented failure taxonomy.

Stage 2: Internal workflow automation (months 6–18, overlapping)

Automate high-volume internal processes with clear rules and measurable accuracy: security alert triage (Oklahoma SOC model), procurement document review, HR inquiry handling, compliance monitoring. These workflows have clear right answers, measurable error rates, and recoverable failures. They generate the performance data needed to calibrate autonomy thresholds for citizen-facing contexts. Success criteria: task completion rate within defined parameters, exception rate within expected range, audit trail passing internal review.

Stage 3: Citizen-adjacent support (months 12–24)

Deploy agents in support roles for citizen-facing processes — information gathering, document completeness checking, status communication — while human staff retain decision authority. Agents prepare and support; credentialed officials decide and communicate. This stage builds public trust infrastructure: transparency mechanisms, feedback channels, and the communication strategy for explaining AI involvement in government services. Success criteria: citizen satisfaction scores, escalation rate within expected range, successful FOIA response for agent decision rationale.

Stage 4: Supervised citizen-facing applications (months 18–36)

Deploy agents in roles with limited direct citizen impact: scheduling, FAQ response, status inquiry handling, application completeness guidance. Maintain robust escalation to human staff. Governance requirements are highest here: accessibility compliance (Section 508), explicit human review before any consequential determination, and audit trails capable of satisfying congressional or judicial inquiry. Success criteria: accessibility audit passing, zero unreviewed consequential determinations, public reporting on agent usage and performance.

Platform and Procurement Considerations

Government procurement for agentic AI platforms requires questions that standard commercial technology procurement does not adequately address.

FedRAMP authorization as a baseline, not a differentiator. ServiceNow AI Agents’ FedRAMP High Authorization at Impact Level 5 demonstrates that enterprise agentic platforms can meet federal security standards. For agencies handling sensitive data, FedRAMP High authorization — or equivalent DoD IL4/IL5 authorization for defense contexts — should be a procurement prerequisite, verified through the FedRAMP Marketplace, not vendor self-attestation. Provisional Authorization to Operate (P-ATO) status is not equivalent to full authorization for production workloads.

OMB M-25-22 compliance architecture. Before contract execution, require vendors to demonstrate — at the infrastructure level, not the contractual level — that operational data does not flow to model training pipelines. This means: network architecture diagrams showing data isolation, contractual prohibition on training-data use with specific technical enforcement mechanisms, and audit rights to verify compliance. Contract language alone is insufficient; the architecture must enforce the requirement.

Model explainability for procurement contexts. Agencies subject to FOIA and congressional oversight need vendors who can produce decision rationale at the application layer — not just at the model layer. Ask vendors directly: can your system produce a plain-language explanation of why a specific agent recommendation was made, in response to a specific FOIA request, without requiring access to model internals? If the answer requires model interpretability tooling that is not part of the standard deployment, the system does not meet government explainability requirements.

Vendor data handling practices. Beyond OMB M-25-22 compliance, government agencies should require full disclosure of: subprocessors who may access operational data, data residency (US-only storage for sensitive workloads), incident notification procedures that meet federal requirements, and exit provisions that ensure data portability and deletion on contract termination.

Building the Governance Capability

The 2025 OMB memos that formalized Chief AI Officer authority created an accountability structure. The question for each agency is whether the governance capability exists to exercise that authority effectively.

Chief AI Officer first-year priorities:

The Chief AI Officer’s first priority is not deployment — it is inventory. Before building new agentic AI capabilities, the CAIO needs to know what AI systems are already operating in the agency: which systems use AI components, what data they access, what decisions they influence, and what oversight exists. Many agencies will discover AI capabilities already embedded in enterprise software they have been using for years, without governance structures in place.

The second priority is policy. OMB M-25-22 requires agencies to develop AI use policies. That policy must address, at minimum: the classification of AI use cases by risk tier, the approval process for new agentic AI deployments, the human oversight requirements by risk tier, and the incident reporting and response procedures. The policy should be developed with legal, compliance, and mission owners — not drafted by the CAIO’s office and handed down.

The third priority is capacity. The governance infrastructure only functions if the people operating it have the skills to do so. This means: training for program managers on AI risk assessment, training for contracting officers on AI procurement requirements, training for supervisors on how to evaluate agent performance, and a designated AI review capability — internal or contracted — that can conduct technical assessments of proposed deployments.

The governance committee:

Establish an AI Governance Committee that meets at minimum quarterly, with representation from: the CAIO’s office, legal/general counsel, CIO, privacy officer, mission program leads, and inspector general or audit function. This committee reviews new deployments before authorization, receives incident reports, reviews performance data, and updates policy as the regulatory environment evolves. Its existence — and its records — is what demonstrates to oversight bodies that governance is not notional.


Make It Your Own

Key questions to ask in the context of your organization:

  • Have you identified which of your agency’s workflows are appropriate for high-autonomy agent operation versus those requiring human sign-off, and documented that distinction formally in your AI governance policy?
  • Does your vendor contract and system architecture explicitly prohibit the use of non-public government data for model training, as required by OMB M-25-22?
  • Have you designated a Chief AI Officer or equivalent with defined authority over agentic AI programs, and does that person have the resources and organizational standing to enforce standards?
  • Can your agentic systems produce plain-language decision rationale that would satisfy a FOIA request or congressional inquiry without requiring technical translation?
  • Have you built and tested rollback procedures for your highest-volume agentic workflows before moving them to production?
  • What is your agency’s sequence for building from internal-facing agent applications to citizen-facing ones, and what governance milestones must be met before advancing to each stage?