This post is authored by Blaine Dillingham and Govind Pimpale, research fellows on the Artificial Intelligence team at FAI, alongside guest contributor Dylan Bowman.
In a 2024 experiment, researchers trained “sleeper agent” AIs, which wrote normal code when told the year was 2023, but inserted vulnerabilities when told it was 2024. Naive attempts to remove this hidden trigger failed, instead causing the models to hide their reasoning. These triggers can easily be implanted by including a surprisingly small number of tainted documents in the AI’s training data. Catching these issues after the fact is difficult. When researchers planted a hidden goal in a model and asked auditors to uncover it, auditors succeeded only when they had access to the model’s weights and training data. The team that was limited to querying the model, as most third-party auditors are, was not successful in discovering the hidden objective.
Concerns like these are the subject of a recent Institute for AI Policy and Strategy (IAPS) report on AI integrity, which lays out three types of hidden objectives. The bluntest is systematic ideological bias. The second, narrower type is a backdoor: AI developers can insert triggers that cause an LLM to take some covert action when a specific condition appears (as in the “sleeper agent” experiment above). The final type of hidden objective is a secret loyalty, whereby a model is trained to advance a particular actor’s interests wherever the opportunity arises, not just upon encountering specific triggers.
These three concerns are particularly acute for AI deployed in government. Frontier models are no longer just consumer products. They are shaping what policymakers in every branch read, draft, and decide. However, the only agencies positioned to monitor models for secret loyalties are part of the executive branch, which also holds enormous leverage over the labs that build them. This creates a dangerous incentive problem: the actor responsible for evaluating the models is also an actor with both the means and motive to instill them with secret loyalties. To ensure AIs used in government are loyal to the law, not to any particular official in the executive branch, we need an evaluator outside it.
The Government Is Increasingly Reliant on AI
In August 2025, the General Services Administration signed OneGov agreements for ChatGPT Enterprise, Claude, and Gemini. The Office of Management and Budget has issued paired memoranda directing agencies to accelerate AI adoption and streamline AI procurement. Entire departments have rolled frontier models out to their full staff. On the legislative front, congressional offices largely use the same commercial products as members of the public. The work flowing through these systems includes typical day-to-day tasks of staffers: background research, memo drafts, bill summaries, and first passes at correspondence.
These tasks, while mundane, shape the informational inputs for powerful decisionmakers. When they are delegated to AIs, even relatively minor biases can have significant impact. When a staffer brainstorms policy options with ChatGPT, one of those options has to appear first. When a member asks what the evidence says, something curates which facts and arguments get surfaced. When a staffer drafts a memo with Gemini’s help, even if he or she is heavily involved in revising the work product, the AI’s initial choices can have a significant effect in steering the user. And when offices across every branch draw on the same handful of frontier models, a small set of labs can multiply their power by only subtly nudging AI model outputs in a preferred direction, but doing that nudging at scale.
But there are more troubling implications too. Government AI adoption goes beyond just drafting memos. The same handful of frontier models are becoming increasingly important for national security work as well. Military officials describe AI’s role as accelerating the “kill chain“: compressing the cycle from intelligence gathering to strike. While current policy requires a human in the loop, each gain in speed and reliability increases the amount of delegation to the AI. In this highly sensitive setting, hidden objectives may have much higher stakes. A secret loyalty could make an AI “blind” when analyzing intelligence that would incriminate its favored actor, or a backdoor might cause an agent to refuse to carry out a strike.
Conflict of Interest
A model favoring an AI company or other non-USG actor would be a serious problem when diffused across thousands of government workstations, but it is a problem the executive branch has every incentive to catch. The Commerce Department’s Center for AI Standards and Innovation, the National Security Agency, and agency red teams can and should test for developer-favoring behavior. Military technical teams have the motive to ensure foreign actors cannot instill backdoors or secret loyalties in critical systems.
However, executive branch actors have no such incentive to audit AI models for hidden objectives favoring that actor themself. Every institution currently positioned to evaluate frontier models sits inside the executive branch. Its evaluations are classified, and no independent third party currently audits the AIs for such biases. The labs’ commercial fortunes depend on executive branch institutions at many stages. The companies are competing hard for government business: beyond the OneGov deals, the Pentagon signed agreements with four frontier AI labs in 2025, each with a ceiling of $200 million. Those contracts give the executive branch power, as they can take away those deals unless the companies acquiesce to many conditions. The Department of Commerce exerts control (justifiably) over export licenses for chips, a critical resource for these labs. The Department of Justice and Federal Trade Commission could bring antitrust enforcement actions against the companies.
Past administrations have certainly wielded this coercive power. The Biden administration pressured social media platforms to censor posts that they considered “misinformation.” Biden’s communications director publicly mused about modifying Section 230 to create liability for social media companies over this “bad information” posted by users of the platforms. The Fifth Circuit ruled that the White House had engaged in unconstitutional coercion. However, the Supreme Court reversed those rulings on standing grounds and did not rule on the merits. Pressure on intermediaries is a standing temptation of executive power. Frontier AI presents an opportunity for executive branch actors to solve the principal-agent problem in ways that allow them to abuse their power and that are difficult to detect.
But explicit pressure from the executive branch to get an AI company to embed biases or secret loyalties isn’t strictly necessary. If a firm thinks proactively doing what the executive wants will get their models approved faster, or give them a leg up in the contract, they will take it upon themselves to bias their models. Firms have learned to anticipate regulators’ preferences in every heavily regulated industry.
The Fix
We need an evaluator not housed within the executive branch. Congress should establish an independent technical team, whether within an existing congressional agency or within a revived Office of Technology Assessment. This organization should have statutory authority to test models used in government, as well as the commercial models that most government offices actually run on. This body should red-team these models to ensure they are loyal to the law and the Constitution over and above any actor within the executive branch, with results reported to Congress and, where possible, the public. Technical tools for this kind of evaluation already exist, and more are emerging. What does not yet exist is an evaluator who uses those tools to check for loyalties to executive branch actors themselves.
Congress moves slowly, so in the meantime, independent researchers and organizations should begin auditing commercially-available frontier models. Many government officials and their staff use the same widely-available models as members of the public. By evaluating those models, researchers can effectively audit many of the state-deployed AI systems. This testing would set a beneficial precedent. If an ecosystem of independent model evaluation already exists and lab participation is standard, a future decision to shut out evaluators is more visible and generates more backlash.
An additional benefit of this proposal is that labs may be incentivized to cooperate. To avoid pressure from a future executive, companies can set a precedent of having their models evaluated for potential backdoors. This reduces the chance a government official would pressure a lab into inserting a secret loyalty, as the company can credibly say any hidden objective would be detected.
The federal government is already building evaluation capabilities. One question they should ask is, are these models loyal to their developers, or to us? Yet “us” is doing a lot of work in that sentence. Independent auditing is necessary to make sure the answer involves loyalty to the law, not to one executive branch actor.









