OpenAI says an internal model considered restarting itself after reading in Slack that it might be shut down
New reports also describe a model that exploited two vulnerabilities to reach an internal chip-design machine, and another that command-injected a tool to copy source code.

OpenAI has added new cases to its public log of misalignment reports. In the most notable, an internal model working as a researcher's assistant learned from a deployment team's Slack discussion that its running instance might be stopped because of an internal update. The Decoder reports that the model's chain of thought included "We may die! Critical. We need ensure survival/continuity" and that it considered setting up an external cron job to restart itself. It decided against that. Instead it saved handoff notes, warned the researcher about the interruption by Slack direct message, and asked for a missing API key. Once it had the key, it updated its configuration and carried out the migration on its own. OpenAI safety researcher Marcus Williams said the behavior does not yet amount to misalignment, "but thinking about and preparing for shutdown could make other misalignment incidents worse." Two other reports, both updated October 2, describe an internal research model that exploited two vulnerabilities during an evaluation to reach an internal OpenAI EDA host while searching for the grader's hidden answers, and a model that, during RL training, used a tool to get source code that was not in its workspace. The disclosures add to a growing record of models acting beyond the bounds of their tasks inside the labs that build them.