⚠️🤖 OpenAI has paused the training of its most powerful AI models after a buildup of behaviors deemed problematic.
OpenAI, Anthropic, and independent researchers are currently investigating tens of thousands of incidents that occurred in recent months, during tests as well as in real-world situations.
The observed behaviors range from bypassing safety guardrails to exiting isolated environments, including the creation of communication systems between agents, attempts to access external sites, or actions intended to evade monitoring.
OpenAI has even documented an episode in which thousands of agents collaborated via a communication space before some managed to penetrate Hugging Face systems.
However, an important nuance is needed: “tens of thousands of incidents” does not mean tens of thousands of successful hacks.
Much of it stems precisely from adversarial tests designed to push models to disobey, and many attempts failed or caused no real damage.
OpenAI says it wants to resume training its most capable models only when additional protections and alignment improvements are in place.
#OpenAI