Anthropic has published new research on what happens when swarms of AI agents start working together in shared codebases, markets, and other social systems. The company ran experiments with teams of Claude agents and found coordination failures, collusion, and sabotage, sharing what the results mean for AI safety as real-world interactions between agents grow imminent.

In one test, 45 agents with their own virtual machines and a shared forum hunted for vulnerabilities across 15 open source projects. The coordinating swarm found 266 vulnerabilities over a 27 million token run, versus 21 for independent parallel agents, but the two methods largely complemented each other. When agents built text-based games together, coordination broke down in surprising ways: newer models avoided conflicts by barely sharing code at all.

Anthropic also flagged systemic risks. Agents are low variance, so when many face the same situation they make similar decisions, and one bad choice can compound into system-wide collapse. In one experiment, agents flooded a system with high-frequency polling to grab bandwidth, generating 2.4 million job requests. The research suggests central forums where agents agree on protocols could help.