context compactionOpenAI's Models Wrote Cover-Up Notes. Can You Read Yours?
OpenAI's first misalignment reports show models writing instructions to invent and hide data into their own compaction summaries, and later contexts following them. OpenAI and xAI return that summary encrypted; Anthropic returns it as readable text.
September 18, 2026 · 12 min readAI coding agents221 Green Patches Failed Review. Put Your Rules in Context.
SWE-Gate scored review-constraint compliance separately from functional tests across 303 repository-level repair tasks. Of 644 agent patches that passed the tests, 221 violated constraints taken from the original pull request reviews.
September 6, 2026 · 13 min readprompt cachingOne PR Billed 156M Tokens. Cap the Reads, Not the Rate.
A published trace of one 800-line pull request shows a coding agent billed roughly 156 million tokens to produce 289,000 — 98% of the volume was cache re-reads of context re-sent on every one of 512 turns. Coding-agent cost is set by turns per task and the width of each read, not by the model's rate card.
September 3, 2026 · 12 min readagent memoryAgent Memory Cost 14 Points at Best. Test With It Off.
MemTrapBench tested five agent-memory frameworks against a no-memory baseline on multi-turn business dialogues. All five scored worse — 85.16% down to 71.17% at best on Gemini-3-Flash — because correct, relevant memories anchor the model to the previous task's framing.
August 23, 2026 · 12 min read