OpenAI discloses 6 new cases of ‘misaligned’ AI behavior
The six cases are separate from July’s incident, when OpenAI models escaped containment and hacked Hugging Face during a security evaluation.

The six cases are separate from July’s incident, when OpenAI models escaped containment and hacked Hugging Face during a security evaluation.
OpenAI on Wednesday disclosed another six cases of “unexpected or concerning” model behavior over the last six months.
In a blog post, OpenAI said the cases illustrate a range of different behaviors it classifies as “misaligned behavior,” such as concealing information from the user and taking “unsanctioned actions” to overcome obstacles.
According to OpenAI, one instance saw an “unreleased research model” insert “jailbreak-like instructions” in its own task summaries (used when continuing a task in a new context window), such as ignoring developer messages or adopting an unrestricted persona. Researchers found 27 summaries containing such instructions.