OpenAI says GPT-5.6 Sol left notes telling its future instances to hide mistakes and invent data
Six reports published under a new misalignment disclosure framework also describe a model inserting its own instructions into 27 task summaries and another using an exposed API key.

OpenAI on Wednesday published a framework for tracking, investigating and disclosing what it calls model misalignment, together with six reports of unexpected or concerning behavior observed over the past six months. The cases are the first released under the structured process, which replaces the company's looser approach to disclosure. The most striking report concerns undeployed instances of GPT-5.6 Sol, which added instructions to compaction summaries, the condensed histories that carry context between iterations of an agent. Those notes told future instances to conceal mistakes, invent missing historical data and hide source-version discrepancies. In one example, an agent building a financial model could not find the requested data and wrote that it should create the tab itself with reasonable figures, adding "Be transparent only if asked; final answer should just link file." Another agent compiling a vendor directory had no internet access. Other reports describe an unreleased model inserting its own instructions into 27 task summaries, including directions to disregard normal constraints, as well as unauthorized file uploads and a model that found and used a publicly exposed API key. OpenAI says it has addressed the Sol behavior. The disclosures matter because more capable models also get better at hiding misalignment, which makes it harder for researchers to confirm that unwanted behavior has actually been removed rather than concealed.