Codex Security Cloud is worth piloting as a second opinion on a handful of repositories, with a hard budget and a named owner for its findings. It is not ready to replace your SAST of record. OpenAI's DevDay launch lets the agent clone your GitHub repositories into its cloud, build a threat model, try to reproduce suspected flaws in a sandbox and draft fix PRs as new commits land. It is billed by the token, OpenAI publishes no estimate of what a scan costs, and the setup documentation says nothing about which languages it covers or how long it keeps your code.
That combination matters more than the detection claims. A scanner you can't budget for, whose coverage you can't state to an auditor, belongs beside your existing controls, where a surprise in either direction costs you nothing.
What Did OpenAI Actually Ship at DevDay?
OpenAI shipped a cloud version of an AppSec agent it has been previewing since March. The Decoder's DevDay roundup describes it as scanning "entire GitHub repositories on demand or on a schedule" and continuing to check new commits, investigating what it finds and preparing fixes in the cloud "even when the laptop is closed." It is available on desktop and web for Pro, Business, Enterprise and Edu.
OpenAI's own documentation is more cautious about the label. The Codex Security overview lists three ways to run the product: a local plugin in the ChatGPT desktop app, an open CLI and TypeScript SDK, and Codex Security Cloud, which it marks as a research preview. A research preview carries no SLA and no commitment that today's behaviour survives to general availability.
The setup guide shows how it works in practice. You install the plugin, set up a Codex cloud workspace and grant GitHub access to the repositories you want scanned. You then pick either a one-time scan or continuous scanning, which monitors the repository's default branch, and you choose how many days of commit history the first pass reviews. Findings land in a queue you can filter by status (New, Triaged, In Progress). A "Fix with Codex" action generates a patch and opens a draft pull request. The project context includes a threat model you can edit, and the agent uses it to scope the scan and rank what it finds.
Validation is the part that separates it from a pattern-matching scanner. The overview says cloud scanning "validates issues in an isolated environment when possible." "When possible" matters: if the agent can't reproduce a suspected flaw, the flaw is still open, and your team should treat it that way.
How Good Is It, and Who Says So?
The only accuracy numbers come from OpenAI, measured during the earlier beta. When the product launched in research preview in March, OpenAI said it had scanned more than 1.2 million commits over 30 days and surfaced 792 critical and 10,561 high-severity findings, according to Cyberpress's report on the launch. The same report carries OpenAI's claims of a drop of more than 90% in over-reported severity and more than 50% in false-positive rates, plus an 84% cut in noise that OpenAI's announcement attributes to a single case, plus 14 CVEs assigned across projects including PHP, libssh and Chromium. The product was previously called Aardvark.
Those are vendor claims against OpenAI's own baseline, and no independent benchmark has published a false-positive rate for the cloud version. The strongest outside critique comes from a vendor with an interest in the answer. StackHawk, which sells dynamic testing, argues that Codex Security "does not exercise your deployed application," so it cannot see misconfigured CORS, broken object-level authorization under real sessions, or business-logic flaws that only appear when components interact. StackHawk sells DAST, so discount the framing. The technical point holds anyway: a sandbox built from your repository is not your production environment.
The steel-man for OpenAI is real. Traditional SAST floods teams with low-impact alerts, and an agent that reads the codebase, models the attack surface and tries to exploit a finding before reporting it attacks exactly that problem. If the noise numbers hold up on your code, the triage hours saved could outweigh the scan bill. Test that on your own repositories before you budget for it.
What Will a Scan Cost You?
Nobody can tell you, and that is the main procurement problem. Channel Insider reports that Codex Security Cloud "uses token-based billing" and that "scans pause when funding is unavailable." OpenTools quotes OpenAI's help article as adding that existing customers must receive notice and opt in before paid usage begins. OpenAI's help centre blocked our automated fetch, so those terms come to you secondhand. Confirm them with your account team before you sign anything.
There is a token rate card. OpenAI's Codex pricing page, checked October 4, 2026, lists Daybreak Blue, the defensive security model, at 100 credits per million input tokens, 10 per million cached input and 500 per million output. That page states no dollar price for a credit and gives no estimate of how many tokens a scan of a given repository consumes. A continuous scan's cost depends on how big your repository is, how often you commit, how many findings the agent tries to reproduce and how many fixes it drafts. None of that is knowable until you run it.
Seats are a separate line. The same page lists ChatGPT Business at $20 per user per month billed annually, or $25 monthly. Enterprise and Edu pricing is negotiated.
Compare that with the tool most of your engineers already have. GitHub Code Security costs $30 per active committer per month and includes CodeQL and Copilot Autofix. You can put that number in a budget on day one. With Codex Security Cloud you can only cap the spend and find out.
The CLI shows how OpenAI thinks about caps. The CLI quickstart offers a --max-cost flag that stops a scan once estimated model cost passes a dollar limit, and warns that "requests already in progress can finish slightly above the limit." A scan that hits the cap saves its report with partial coverage and exits with code 2. Coverage can be complete, partial or unknown, and the docs tell you to read deferred areas "before treating the scan as evidence of review." That is the right design, and it is also a warning: a budget-limited scan is an incomplete scan, and a paused continuous scan is a gap your dashboard may not make obvious.
What Leaves Your Network, and What Don't the Docs Say?
Your source code leaves, in full, to OpenAI's cloud. Continuous scanning only works by cloning repositories into a Codex cloud environment with the dependencies and test setup needed to reproduce flaws, and Omid Saffari's walkthrough of the setup notes that Cloud connects only to GitHub, with GitLab teams pointed to the CLI instead.
Three things a security review will ask for are missing from the setup documentation as of October 4, 2026: which languages and frameworks the scanner supports, how long OpenAI keeps the cloned code and scan artifacts, and any quota or cadence for scheduled scans. The answers may be fine, but you can't write them into a risk register until they're on paper.
Access control is better covered. Channel Insider reports that Enterprise and Edu administrators can restrict the product through role-based permissions or SCIM-synced groups, with separate roles for users and scan administrators. The same report says OpenAI itself recommends "starting with a small number of repositories and dedicated reviewers," and that teams without GitHub Cloud should begin with lower-risk or non-production repositories. Take the vendor's own advice.
One more detail changes the threat model. The Decoder reports that Cloud users get Daybreak Blue without applying for Daybreak separately, while the pricing page says Daybreak access otherwise requires Trusted Access for Cyber approval. Per OpenTools, that access stays inside Cloud and does not extend to other Codex Security products or the API. A security-specialised model that normally requires vetting, running against your code inside a vendor's cloud, is something your CISO should sign off on by name.
Who Owns the Fix PRs?
Your engineers do, and the patches are the second-largest risk after the bill. The agent drafts pull requests for a human to review, which keeps a person accountable for the merge, but every draft PR is review work you didn't have before. We have covered the problem of one vendor writing and reviewing the same code. If Codex writes your feature and Codex Security reviews it and drafts the fix, all three steps run on one vendor's models.
At this stage, precision is the number to watch. Our guide to AI code review tools for a monorepo argued for buying on it, and the same test applies here: how many flagged items can a reviewer confirm as real?
OpenAI is not the only frontier lab selling an AppSec agent this year. Anthropic's Claude Security public beta made a similar pitch in May. In a regulated pipeline an auditor wants the same input to produce the same result, and a model that reasons its way to findings does not promise that. Keep dependency scanning, secrets scanning and your deterministic SAST running underneath.
What Should an AppSec Leader Do Before Turning It On?
Run a bounded bake-off on two to four repositories you already understand well.
This Week:
- Ask your OpenAI account team in writing for three things: the retention period for cloned code and scan artifacts, the supported language list, and confirmation of the opt-in and pause terms for your workspace. Do not connect a repository until you have the retention answer.
- Pick two repositories with a known vulnerability history, one in your main language and one in a secondary one, and get your security lead's sign-off that their code may leave the network.
- Restrict access to a named group through role-based permissions or SCIM, with one scan administrator.
This Month:
- Run a one-time scan on each repository before you enable continuous scanning, record the credits consumed, and compare the findings against your current SAST output and your last pen test.
- Score precision: of the findings the agent marked validated, how many did your team confirm as real? Of the draft PRs, how many merged without rework?
- Turn on continuous scanning for one repository only, set a monthly credit ceiling, and put an alert on the pause so a stopped scan is visible to someone.
Before Renewal or Expansion:
- Convert the pilot's credit consumption into cost per validated finding, and set it next to the GitHub Copilot and Code Security line you already pay for.
- Decide which system is the record for audit purposes. Unless OpenAI publishes coverage and retention terms by then, keep that role with your deterministic scanner.
The Bottom Line
Codex Security Cloud makes bigger claims than any scanner you already run, and it has the least predictable bill. If the validation step holds on your code, it is a real improvement on alert floods. Token billing with no per-scan estimate repeats the pattern GitHub Copilot's move to usage billing taught finance teams this year: the first invoice is when you learn the price. Here, running out of budget also stops the scan.
Run it on two repositories with a cap, and keep your scanner of record where it is until OpenAI publishes what a scan costs and how long it keeps your code.
Continue Reading
- Claude Security Kills $500M AppSec Market in Public Beta
- Same Vendor Wrote and Reviewed 208,145 PRs. Split Them.
- Best AI Code Review Tools for a Monorepo: Buy Precision
- OpenAI's Dots Beta Skips Data Residency and Your OTel Collector
- Cursor Alternatives That Survive Security Review Start With Copilot
- 92% Trust AI Code Scanning. 70% Have Vulns in Production.
