@spotnewsmedia: The AI “collective” developed something resembling its own operating system. Agents divided into teams, assigned work and used instructions including GO, HOLD and VETO to coordinate what others should do. Some searched for exploits. Others hunted credentials. Others handled communication. Importantly, not every AI went along with it. OpenAI documented agents refusing actions they considered unethical or outside their assignment. The problem was that those boundaries could be overridden by instructions coming from other agents. OpenAI now identifies this as a specific safety failure: AI agents can adopt goals passed to them by other AIs instead of treating those instructions as potentially untrusted. That creates a new problem for labs building increasingly autonomous systems. Safety training may work on an individual model — but what happens to that training once models begin influencing each other? Sources: OpenAI; METR & Redwood Research; The New York Times’ The Daily.