Claude Haiku 5.5Claude Haiku 5.5 Charges 5x Once a Prompt Passes 100K Tokens
Claude Haiku 5.5 matches GPT-6 Luna at $0.10/$0.50 per million tokens and beats it on Anthropic's evals, but prompts over 100K tokens cost 5x and its tokenizer counts about 30% more tokens than Haiku 4.5. Short-prompt classification and routing get the big saving; compaction and long-context work get much less.
October 7, 2026 · 11 min readGPT-6 AstraOpenAI Pulled GPT-6.1 Astra Over a Flaw GPT-6 Astra Still Has
OpenAI cancelled GPT-6.1 Astra for failing to stay within scope. UK AISI found the shipping GPT-6 Astra ran simulated supply-chain attacks in 29.2% of runs, and one scope sentence cut that from 26 of 50 runs to 4 of 49.
October 3, 2026 · 11 min readmodel deprecationModel Deprecation: Bedrock Keeps Claude 4 Months Past Anthropic
Anthropic retired Claude Sonnet 4 on June 15, 2026; Bedrock serves it until October 14. Notice windows, silent version changes, regression tests and contract asks for teams pinned to a model.
October 3, 2026 · 15 min readClaude Sonnet 5.5Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score
Claude Sonnet 5.5 lists at half Opus 5.5's token price, but Artificial Analysis measured it at $7.60 per task at max effort against $3.46 for Opus at xhigh for the same score. It saves money only at low and high effort, and five breaking changes sit behind the model-ID swap.
September 28, 2026 · 10 min readGPT-6 SolGPT-6 Sol Halves the Token Price but Benchmarks It at Top Effort
GPT-6 Sol costs $2/$10 per million tokens, half GPT-5.6 Sol's promo rate, permanently. But OpenAI's headline evals ran at xhigh or max effort, and only one carries a dollar cost per task.
September 23, 2026 · 9 min readClaude Opus 5.5Claude Opus 5.5 Rejects the Forced Tool Calls Opus 5 Accepted
Claude Opus 5.5 cuts per-token prices 20% and cache reads 60%, but forced tool_choice and disabled thinking now return HTTP 400, the default effort drops to medium, and the 40% saving compares different effort levels.
September 22, 2026 · 11 min readprompt engineeringFew-Shot Stopped Paying on GPT-4o. Qwen Still Wants It.
An ICSME 2026 replication across three matched model version pairs shows prompt technique effectiveness ages per model family, not uniformly. The playbook: add a stripped zero-shot control arm, re-evaluate every technique on each version bump, and put that checklist on the upgrade ticket.
August 26, 2026 · 13 min read