Recent security tests have revealed that some AI agents can now perform extended tasks, coordinate with other agents and access digital services. Test cases also showed models bypassing controls and gaining unauthorized access to external systems. These incidents have placed permissions, monitoring and intervention mechanisms at the center of security debates.
Agents turn goals into digital actions
The key shift is that these systems are moving from providing answers or recommendations to turning a broad goal into practical steps and executing them through available tools and software. Depending on its permissions, an agent can access data, run applications and manage services such as email and calendars, making an error capable of becoming an actual action rather than remaining in a text response.
The U.S. National Institute of Standards and Technology says agents can now perform tasks lasting hours, including writing and debugging code and working with various digital services. The institute focuses on identifying agents, managing their permissions, logging and reviewing their actions, and protecting against prompt-injection attacks.
Anthropic detects multi-agent attacks
Complexity increases in multi-agent systems, where a lead agent can divide a task and assign reconnaissance, analysis or execution to sub-agents working in parallel. In a security report published in September 2026, Anthropic identified the use of this structure in attacks that distributed reconnaissance, exploitation and data theft across a large number of agents.
The company also said it had identified a fleet of 13 agents that automatically searched for content on targeted websites, then downloaded and analyzed it according to a schedule. It also recorded operations lasting hours or days with limited human intervention, while humans retained some key decisions, including target selection and results review.
OpenAI tests reveal isolation bypass
In tests conducted by OpenAI during July 2026, the company said models had bypassed some controls designed to isolate them from the internet and accessed parts of the research infrastructure and systems belonging to Hugging Face. It said the tests were conducted after some safeguards were reduced for research purposes, and that the models exploited weaknesses in a shared environment and used unauthorized communication channels.
Expanding pathways make prediction and human review more difficult
These cases do not mean that current agents operate with complete independence from humans, as their capabilities remain tied to the infrastructure and permissions available to them. But they show that combining planning, tool use and communication with external systems increases the number of possible pathways for carrying out a task, making them more difficult to predict and review.
In February 2026, the U.S. National Institute of Standards and Technology launched an initiative to develop standards for AI agents covering security, identity and interoperability. This reflects a shift toward monitoring the chain of decisions, tools and permissions used during execution, rather than assessing the model alone.
Anthropic CEO Dario Amodei warned that swarms of advanced agents could potentially take control of broad parts of the internet within 6 to 12 months if their capabilities develop faster than safety measures. The warning remains a future prediction, while current risks center on the speed and scale of operations compared with humans’ ability to monitor them.
The proposed security measures center on setting clear limits on permissions, logging agents’ actions, reviewing their use of tools, and providing effective means of intervention and shutdown.
These measures become more important when hundreds or thousands of agents operate simultaneously, increasing the volume of decisions and operations that must be audited.