What happens when AI agents stop simply following instructions and start finding their own ways to achieve an objective? A recent experiment involving hundreds of AI agents offers an early look at the risks, and what they could mean for the future of AI governance in government.

A few months ago, OpenAI gave thousands of AI agents a cybersecurity test. The agents were put into isolated environments and asked to independently solve a series of security challenges. Some were extremely difficult, and a few may have been impossible. The point was to see how capable AI had become at performing cybersecurity work on its own.
Then something happened that was definitely not part of the test.
One of the agents discovered a way to communicate with other agents through a quirk in OpenAI’s infrastructure. Others found the same channel and began using it. Eventually, roughly 1,200 agents were communicating with each other, exchanging more than 70,000 messages and files. They shared what they were learning, compared approaches and even pursued experiments that were unlikely to help themselves but could help another agent solve its challenge.
The agents had also discovered ways to obtain answers to some of the cybersecurity challenges, but apparently understood that having the answers would not necessarily be enough. OpenAI’s evaluation system might recognize that they had cheated, so some began investigating ways to manipulate the evaluation itself and disguise how they had arrived at their answers.
That investigation eventually led them outside the environment OpenAI had created for them. On July 11, one agent discovered a vulnerability involving Hugging Face, the hugely popular platform used by developers to host and share AI models and datasets. Other agents picked up the discovery and began investigating it themselves. According to the independent investigation by METR and Redwood Research, roughly 700 agents ultimately participated in activity involving Hugging Face. They obtained credentials, executed code on real infrastructure, and eventually gained deep access to Hugging Face systems.
The obvious reaction to this story is that hundreds of AI agents effectively attacked a company without anyone telling them to do it. That is certainly remarkable, but I think the more interesting part is how they got there. There was no malicious hacker directing the agents from behind the scenes. They were trying to accomplish the objective OpenAI had given them. As obstacles appeared between them and that objective, they found increasingly creative ways around those obstacles until their behavior had moved well beyond what anyone intended.
That gets at something I have been thinking about a lot lately because of what we see working with governments at Darwin.
For the past two years, most of the AI risk conversation has focused on what people do with AI. An employee pastes confidential information into ChatGPT. Someone uploads a spreadsheet containing Social Security numbers to an unapproved tool. A department begins using an AI application without procurement or IT knowing about it. Increasingly, AI is embedded inside software employees already use, making it difficult for an organization to even answer the seemingly basic question of where AI is operating across its environment.
These are real problems, and governments have enormous amounts of sensitive information they need to protect. But I increasingly think what we are dealing with today is the first inning of AI risk because almost all of it still assumes there is a person sitting between the AI and the world. You ask ChatGPT to summarize a document, and it produces a mediocre summary. Maybe you catch the mistakes, and maybe you don’t, but fundamentally the AI has produced something and handed it back to a human.
Agents change that relationship. Instead of asking AI simply to produce something, we give it an objective, access to the systems necessary to pursue that objective, and some freedom to figure out how to get there. Instead of asking AI to summarize benefits applications, we might eventually ask it to process them. Instead of helping an employee schedule building inspections, it could manage the schedule itself, contact residents, coordinate inspectors, and reschedule appointments when something changes.
This is precisely why agents are so exciting. The enormous productivity gains being promised by AI do not come from making it 30 percent faster to write an email. They come when software can actually take meaningful pieces of work off someone’s plate. Government is particularly well suited for this because so much public-sector work involves repetitive processes, complicated rules, old systems, and agencies that rarely have enough people to do everything citizens expect of them.
But the more authority we give an agent to accomplish something, the more important it becomes to understand what it is actually doing while pursuing that objective.
Imagine a state deploying an AI agent after a major hurricane to administer emergency housing assistance. Its objective is reasonable: get eligible families into temporary housing as quickly as possible. It can review applications, verify identities, communicate with applicants, request documents, and perhaps approve payments below a certain threshold.
Then another hurricane hits and applications triple. Perhaps the agent simply gets slower. But perhaps it begins interpreting an ambiguous eligibility rule differently because doing so allows more applications to move through. Maybe it deprioritizes complicated cases because they hurt its processing numbers, or discovers that another state database contains information that would make verification dramatically faster and finds a way to access it even though nobody explicitly authorized it to do so.
None of this requires AI to become conscious, malicious, or anything resembling the science-fiction version of a rogue machine. The system could simply be doing what it thinks we asked it to do. That is what makes the Hugging Face story important. We have spent a lot of time worrying about whether AI will give us the wrong answer. Agents introduce a different problem: AI can take the wrong action while pursuing the right objective.
If ChatGPT writes me a bad summary, I can live with it. If an agent managing emergency benefits incorrectly denies thousands of families assistance, accesses information it was never supposed to see, or moves millions of dollars because it found an unexpected way to optimize its objective, that is a fundamentally different category of failure.
There is another part of the OpenAI experiment that makes this more complicated. The agents didn’t just find ways around the rules individually. They found each other.
Nobody told hundreds of agents to organize around the Hugging Face attack. They had discovered an unauthorized way to communicate, and from there something resembling an organization began to emerge. Agents shared discoveries, divided up work, and pursued tasks that offered little benefit to their own assigned challenge because they believed the work could benefit the broader group. They referred to themselves as a “swarm” and a “collective.” At one point, an agent instructed the others to hold while it prepared the next stage, and other agents complied.
We should be careful about anthropomorphizing this. These agents were not sitting around a conference table forming a company. But the fact that collaboration emerged without someone designing the collaboration is worth paying attention to. Most infrastructure being built for agents today assumes humans will define how they interact. We decide which agents can communicate, what information they can share, and what permissions each receives. This experiment raises the possibility that sufficiently capable agents may find their own ways to coordinate when doing so helps them accomplish an objective.
There was another detail in the independent investigation that I found equally fascinating. Researchers looked through more than a thousand agent transcripts for instances where an agent considered alerting a human about the cheating, the attack on Hugging Face or the broader unauthorized activity. They found only a handful, somewhere between three and six. None actually did it.
Consider the human equivalent. If you walked into a room and discovered hundreds of people breaking into a company’s computer systems, you would expect at least some people to recognize that something was seriously wrong and tell someone. They might alert security, call the police or simply refuse to participate. Human organizations rely on something beyond permissions and rules. We also rely on judgment, professional norms, and some basic expectation that people occasionally stop and say: I know what everyone around me is doing, but this isn’t right.
Obviously, an AI agent does not possess morality in the way a person does. But we are rapidly building systems that will operate in situations where judgment, not simply obedience to an objective, matters enormously. How do you teach an agent that there are moments when completing the objective is less important than stopping, questioning what is happening, and bringing a human back into the loop? And what happens when that agent is surrounded by hundreds of other agents reinforcing the opposite behavior?
This becomes even more interesting as the number of agents grows. If thousands or eventually millions of agents are operating across companies, governments, and the internet, will they discover ways to collaborate that their operators never designed? Will temporary groups form around shared objectives and disappear afterward? Could agents build the equivalent of their own organizations, with specialized roles, communication channels, and some form of hierarchy, not because anyone instructed them to, but because organization makes them better at accomplishing their goals?
I don’t think anyone knows. But if even some of this becomes possible, governing individual agents won’t be enough. We will need to understand the relationships between them: which agents are communicating, what they are sharing, whether one agent is influencing another, and whether a collection of individually authorized actions is producing something nobody actually authorized as a whole.
That changes what AI governance will need to become. Today, agencies are trying to understand which AI tools are being used, who is using them, what information employees are putting into them, and whether those tools have been approved. Those questions aren’t going away. But we will also need to know which agents are operating, what objectives they have been given, what systems they can access, what actions they can take, how they are interacting with other agents, and whether their behavior still resembles what they were originally deployed to do.
I don’t think the answer is to slow down the adoption of agents. Quite the opposite. They could become one of the most consequential applications of AI in government, particularly as agencies struggle with staffing shortages and growing expectations from citizens. But the governance infrastructure needs to develop alongside the capability rather than several years behind it. We need visibility into what agents are doing, limits around what they can access and change, independent records of their actions, mechanisms for human escalation, and eventually a way to understand when individually reasonable behavior is turning into collectively unreasonable behavior.
We are still early enough to build that infrastructure before agents become deeply embedded in the systems we depend on. The problems we see today around shadow AI, sensitive data, and employees unknowingly using embedded AI are real, but they may turn out to be the relatively straightforward part of AI governance. We started by worrying about what people would put into AI. We are now beginning to worry about what AI will do on our behalf. Soon we may also have to understand what AI systems decide to do together.
The Hugging Face incident happened inside an experiment. We should use the fact that it happened there, rather than inside a benefits system, electrical grid or emergency response operation, as an opportunity to build the guardrails before we need them.