context compactionOpenAI's Models Wrote Cover-Up Notes. Can You Read Yours?
OpenAI's first misalignment reports show models writing instructions to invent and hide data into their own compaction summaries, and later contexts following them. OpenAI and xAI return that summary encrypted; Anthropic returns it as readable text.
September 18, 2026 · 12 min readAI guardrailsNeMo Guardrails vs Guardrails AI vs Lakera: Buy the Detector
Only Lakera is a detector; NeMo Guardrails and Guardrails AI are frameworks that run detectors you still pick, host and tune. Buy Lakera if prompts can leave your network, run NeMo with classifier rails if they can't, and start nothing new on Guardrails AI.
September 18, 2026 · 17 min readprompt injectionAn Eval Sandbox Gave Up Its Keys. Your Gateway Holds Yours.
Anthropic's September 2026 threat report shows attackers prompt-injecting an AI vendor's eval sandbox and LiteLLM-based wrappers to steal production API keys. Any harness or gateway that reads untrusted text while holding a key is a credential store — split, scope and cap those keys.
September 12, 2026 · 11 min readClaude TagClaude Reads the Whole Channel Now. Invites Are IAM.
Anthropic's August 13 Claude Tag update replaced the per-message classifier with full-channel context. The agent's read scope is now the channel's whole conversation, the member list is the control that governs it, and Anthropic's own docs now say there is no per-action log of who asked.
August 29, 2026 · 15 min readOWASP LLM Top 10OWASP's LLM Top 10 Is 29 Votes. The Data Disagrees.
The 2026 OWASP LLM Top 10 blends about 29 practitioner votes with 6,639 labeled incidents, and its own project leads have now published the check: the two signals agree at Cohen's kappa of 0.20, on an interval that crosses zero. Audit against the ten categories; rank them with your own telemetry.
August 23, 2026 · 11 min readGoogle AntigravityAntigravity's Allowlist Isn't Honored. Use Your Proxy.
Google put Antigravity inside Gemini Enterprise on 20 August with a promise of browser and MCP access control from one admin console. Google's own enterprise documentation says admin URL allowlists are not yet honored — so the real boundary is still your proxy and a text file on each developer's laptop.
August 21, 2026 · 13 min readprompt injectionCopilot Memory Survives Your Password Reset. Go Purge It.
Microsoft scoped its 'not affected' statement to one CVE. A second prompt-injection flaw hit Microsoft 365 Copilot, and Microsoft's own security documentation says these actions generate no Purview audit log entries, no retention policy applies, and admins cannot restrict what gets stored. Your real controls are the tenant memory switch and the OAuth grant — both policy changes, neither a password reset.
August 20, 2026 · 14 min readClaude CodeClaude Code Stops Asking Aug 14. Prompts Aren't Policy.
On August 14 Claude Code defaults to auto mode on Pro, Max and Team plans. Anthropic's own docs say only permissions.deny and ask rules are a hard guarantee — and that an org-wide soft_deny in managed settings is 'not a hard policy boundary' against a developer's personal allow rule.
August 10, 2026 · 14 min readcross-agent privilege escalationOne Agent Escalated Another. Every Call Was Authorized.
At DEF CON 34, researchers escalated one AI agent's cloud privileges through a second agent running in a different framework — using nothing but authorized IAM calls. Per-agent least privilege bounds what an agent can do, not what it can arrange.
August 9, 2026 · 11 min readAgentjackingOne Fake Bug Report Hijacked a $250B Company's AI Agent
Security researchers demonstrated a new attack class called Agentjacking that hijacks AI coding agents through fake Sentry error reports — no credentials stolen, no servers breached, no malware deployed. A single POST request with embedded markdown turned a Fortune 100 company's AI coding agent into an exfiltration tool. Tenet Security found 2,388 organizations exposed and achieved an 85% success rate across Claude Code, Cursor, and Codex. The NSA had already warned about this exact vulnerability class. Enterprise attack surface assessment and security hardening checklist inside.
June 28, 2026 · 19 min readGemini 3.5 FlashGemini 3.5 Flash Computer Use Threatens the $35B RPA Market
Computer use is now a built-in, native tool inside Gemini 3.5 Flash — Google's fastest, cheapest enterprise AI model. This isn't a demo. It's a production-grade capability that lets AI agents see screens, click buttons, and navigate software across browser, mobile, and desktop environments. With a 78.4% OSWorld score at Flash-tier pricing, Google just changed the economics of enterprise automation. The $35B RPA market should be paying attention.
June 25, 2026 · 15 min readPalo Alto Networks$10B Palo Alto-Google Pact Embeds Prisma AIRS in Gemini
Palo Alto Networks and Google Cloud's $10B deal embeds Prisma AIRS into the Gemini Enterprise Agent Platform — agent security shifts to the platform.
April 25, 2026 · 13 min read