@adam_rosler: OpenAI's safeguard for Astra, the first model it rated Critical for cyber, is to read the model's thinking while it works. The Information reports Astra was built with recurrent depth, a technique that loops the same layers inside the model instead of writing every step out as tokens. Cheaper, fewer tokens, and a thinner scratchpad to read. OpenAI wrote last winter that bigger models get more room to think in activations instead of words, and that chain-of-thought monitoring might end up load-bearing anyway. So the monitors they describe now read the activations directly: classifiers on every sampled token, about 20% of the compute being watched, and a flag nobody can clear in thirty minutes pauses the run. Sources: OpenAI, Path to Astra (Sep 1, 2026); OpenAI, Pacing model development in an era of cyber-critical capabilities (Aug 18, 2026); OpenAI, Evaluating chain-of-thought monitorability (Dec 18, 2025); Geiping et al., Scaling up Test-Time Compute with Latent Reasoning, arXiv 2502.05171; The Information, OpenAI Technique in Astra Model Sparks Security Concerns (Sep 1, 2026). Recurrent depth in Astra is single-source reporting that OpenAI has not confirmed.
Adam Rosler
Region: US
Wednesday 02 September 2026 14:01:47 GMT
Music
Download
Comments
hadiauwg16p :
thinking is now actually thinking instead of just text in a special box
2026-09-02 23:03:52
0
david :
haven’t read it yet but seems similar to metas coconut?
2026-09-02 19:17:59
1
hadiauwg16p :
neuralese thinking
2026-09-02 23:03:29
0
To see more videos from user @adam_rosler, please go to the Tikwm
homepage.