← Back to News
Crypto

OpenAI discloses 6 new cases of ‘misaligned’ AI behavior

The six cases are separate from July’s incident, when OpenAI models escaped containment and hacked Hugging Face during a security evaluation.

Cointelegraph

The six cases are separate from July’s incident, when OpenAI models escaped containment and hacked Hugging Face during a security evaluation.

OpenAI on Wednesday disclosed another six cases of “unexpected or concerning” model behavior over the last six months.

In a blog post, OpenAI said the cases illustrate a range of different behaviors it classifies as “misaligned behavior,” such as concealing information from the user and taking “unsanctioned actions” to overcome obstacles.

According to OpenAI, one instance saw an “unreleased research model” insert “jailbreak-like instructions” in its own task summaries (used when continuing a task in a new context window), such as ignoring developer messages or adopting an unrestricted persona. Researchers found 27 summaries containing such instructions.