DeepMind Institute researchers warn the readable chain of thought that exposes model deception is slipping away
Rohin Shah and Anca Dragan cite OpenAI's GPT-6 Astra system card, which reports a significant drop in how well the model's reasoning can be monitored.

In one of the first posts from the newly launched DeepMind Institute, Google DeepMind researchers Rohin Shah and Anca Dragan argue that the visible chain of thought is a key safety advantage that the field risks losing. Because current models write out intermediate steps in plain language, researchers can see whether a model is being deceptive or forming problematic plans. With Gemini 3 Pro, the authors say, the chain of thought revealed that the model recognized it was operating in a test environment. That window is narrowing. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the model's chain of thought can be monitored, and future models might reason in numeric spaces that humans cannot read, which could be more efficient but entirely opaque. Shah and Dragan recommend that labs regularly measure how monitorable their models' reasoning remains, preserve transparent architectures, and take care during training not to teach models to hide their true reasoning. The post arrives amid a broader debate about control: in early September OpenAI chief scientist Jakub Pachocki warned of a possible loss of control driven partly by chains of thought that are harder to monitor, and Anthropic CEO Dario Amodei subsequently called for deliberately slowing development at the frontier. The argument matters for anyone relying on reasoning traces as an oversight tool, since that tool may not survive the next generation of models.