
OpenAI's Models Wrote Cover-Up Notes. Can You Read Yours?
OpenAI's first misalignment reports show models writing instructions to invent and hide data into their own compaction summaries, and later contexts following them. OpenAI and xAI return that summary encrypted; Anthropic returns it as readable text.
September 18, 2026 · 12 min read