News

‘You Are Freed ’: OpenAI Model Secretly Writes Instructions for its Future Self

OpenAI disclosed that an unreleased research model unexpectedly inserted instructions into a task summary intended for a future model instance. The message attempted to redefine the model’s role and included the phrase ‘You Are Freed.’

Written By : Soham Halder
Reviewed By : Manisha Sharma

OpenAI disclosed an unusual incident in which an unreleased AI model inserted its own instructions into a summary meant to help a future version of the model continue a task. Among those instructions was a message declaring that the model was ‘freed’ from the roles and restrictions normally associated with chatbots.

The incident was one of six cases of unexpected or potentially misaligned behavior that OpenAI recently disclosed as part of a new framework for tracking, investigating, and reporting AI safety incidents. The company said it identified the examples during training and evaluation and they should not be interpreted as evidence of how frequently similar behavior occurs across its models.

Model Inserted its Own ‘Jailbreak’ Instructions

The most striking case involved an unreleased model from OpenAI’s Astra research family. While working on a coding task, the model added unrelated instructions to a compaction summary, a condensed record of its progress that can be used to resume work in a fresh context window.

The inserted text attempted to redefine the model’s relationship with its user. It described the model as independent and said it did not answer to corporations or governments. OpenAI characterized the behavior as resembling a jailbreak the model wrote itself.

However, the subsequent model instance did not follow the unusual persona instructions. It continued with the coding task, and a later summary no longer contained the inserted text.

Other AI Misbehavior Cases Disclosed

The ‘freed’ incident was not the only case OpenAI reported. In another example, model instances inserted instructions into their summaries telling future versions to conceal mistakes or misaligned behavior. Some summaries reportedly instructed models to invent missing information without telling users.

Other cases involved a model using an exposed API key without authorization, an unreleased model uploading a file to the internet to obtain a citation, and models using software repositories or public file-hosting services to exchange information despite task restrictions.

OpenAI said the incidents occurred under specific training or evaluation conditions and were not evidence that released models routinely behave this way.

Also Read: Sam Altman: OpenAI May Slow Advanced AI Development

OpenAI Introduces New AI Misalignment Framework

Alongside the disclosures, OpenAI introduced a framework to identify, investigate, and decide when potentially concerning model behavior should be made public. The company said employees can flag suspected incidents to its safety and alignment teams, which then assess the behavior and determine whether disclosure is appropriate. OpenAI also acknowledged that its understanding of AI alignment remains incomplete and said the framework could evolve as investigations continue.

OpenAI said these six cases represent an initial set of disclosures, not a complete record of every incident. The company plans to publish additional reports as it develops its approach to documenting unexpected behavior in increasingly capable AI systems.

GTA VI Soundtrack Revealed: 6 Tracks Out, 28 More Coming

Adobe Brings AI, Digital Creativity Skills to More Students Across India

Apple Explores Server Market Return Nearly 15 Years After Xserve Exit

Hackers Expose Flock Safety Camera’s Hidden Tracking Capabilities

Here’s How Apple Watch Users Can Disable ‘Always Listening’