Astra for Law's headline 54% is real, but it measures how far OpenAI's legal index lifts GPT-6 Astra, not how Astra for Law ranks against other models. OpenAI scored it on 200 questions from the private validation set of Vals AI's Legal Research Bench and compared it only with its own GPT-6 Astra using web search (38.7%). Vals' public leaderboard, which is scored on a separate, never-released test set, already has three general-purpose frontier models at 55.29% with no legal index, and it does not list Astra for Law. Legal ops and law-firm IT buyers should read the launch as "OpenAI closed its own gap," not "the best legal research model." Run a matched test on your own matters before you move any CoCounsel, Harvey or Legora spend.
This is not a knock on the product. A 15-point lift from retrieval is a big effect, and the partner list is serious. But the number in every headline this week answers a narrower question than the one a buyer is asking.
What Did OpenAI Actually Measure?
OpenAI measured one system against a weaker version of itself, on a question set nobody outside can see. Per Legal Technology's launch report, the announcement came on September 17, and the improvement claim rests on 200 U.S. legal research questions from Vals AI's Legal Research Bench private validation set. ChatGPT AI Hub's breakdown gives the figures as a "54.0% all-pass correctness rate on 200 private validation questions" against 38.7% for GPT-6 Astra with web search alone. Both ran at their highest reasoning effort.
Astra for Law is not a new model. MLQ's analysis describes it as a configuration of GPT-6 Astra: the same model, plus a legal search index, legal-specific instructions and confidentiality controls. The planned API model name is gpt-6-astra-law. The index covers more than 230 million URLs of U.S. case law, statutes, regulations, court rules and administrative decisions, and The Next Web reports that a Free Law Project partnership brings in CourtListener's coverage of 99.9% of published U.S. precedential case law.
So the 38.7% → 54.0% comparison isolates exactly one variable: the index. It is a clean test of retrieval, and retrieval won. MLQ notes what it does not test: the results are company-reported and compare Astra for Law with another OpenAI system, not with Westlaw, Lexis+ or any other commercial legal platform.
Why Can't You Put 54% Next to the Leaderboard?
Because Vals publishes leaderboard scores only from its private test set, and OpenAI's number comes from a different set. Vals' methodology page describes three tiers. There is a public validation set for transparency. There is a larger private validation set that Vals "license[s] to companies for their own internal validation". And there is a test set that "remains private at all times, and is the only dataset used for the proprietary benchmark results we publish." Vals says it provides statistical evidence that the private validation set is correlated with the test set. Correlated does not mean interchangeable.
Here is what the public board shows. The Vals Legal Research Bench page was updated September 21. The BenchLM mirror of that leaderboard lists 63 models, with Muse Spark 1.3 Max, Claude Opus 5 and Claude Fable 5.1 tied at 55.29%. Claude Opus 5.5 scores 50.48%. GPT-6 Astra at max reasoning sits at rank 22 with 39.42%. Astra for Law does not appear.
One detail makes the comparison more useful than it first looks. GPT-6 Astra scored 39.42% on the public test set and 38.7% in OpenAI's validation run, within a point of each other. For that model, at least, the two sets behaved alike. That makes it reasonable to read Astra for Law's 54.0% as landing around the frontier cluster rather than above it. It does not make it a head-to-head result.
And even on one shared set, the gap would not be meaningful. At 200 questions and roughly 54% accuracy, the 95% confidence interval is about ±7 points (1.96 × √(0.54 × 0.46 / 200) ≈ 0.069). A 1.3-point gap between 54.0% and 55.29% is well inside the noise. We made the same point about how few model gaps survive a properly sized eval set.
The fair reading: the legal index brings GPT-6 Astra up to roughly where general frontier models already score, not past them.
Is the Index Worth Anything, Then?
Yes. The index is the most defensible part of the launch, but its advantage over rival models is smaller than its advantage over plain web search. Vals' public benchmark harness on GitHub already gives agents web search, page parsing, document retrieval and the CourtListener API for opinions, dockets and citations. So the general models on the leaderboard were not researching blind. They had access to the same case-law corpus that anchors OpenAI's index. OpenAI has not said whether its "web search alone" baseline had CourtListener access, and that choice alone could explain a large share of the 15 points.
OpenAI also reports that Astra for Law found 24% more reference cases on case-law questions and retrieved up to 54% more relevant passages from the correct opinions, per The Next Web. For a litigator, that recall figure may matter more than the pass rate. A missed controlling case is a worse failure than a clumsy paragraph. That is also the claim to test hardest in your own bake-off.
Legal AI Has Run This Test Before
The last independent head-to-head found the general model matched the legal tools on accuracy. The legal tools won on citations. In Vals' VLAIR legal research study, updated October 14, 2025, 200 questions went to Alexi, Counsel Stack, Midpage and ChatGPT. Vals found "little differentiation between the legal AI products (78-81% accuracy, 80% averaged) and the generalist AI product (80% accuracy)." Where the specialists pulled ahead was authoritativeness, where they scored six points higher on average.
That is the pattern to expect again. Accuracy converges fast between specialist and generalist systems. Source quality and citation discipline take longer, and that is exactly the layer OpenAI has now built. The question for your team is whether OpenAI's index beats the one you already pay for. The question is not whether it beats GPT-6 Astra's web search.
It also explains why incumbents are hedging, not panicking. According to Artificial Lawyer, Astra for Law launched with 26 partner-built plugins from firms including Harvey, Legora, Thomson Reuters and iManage. Thomson Reuters' CTO Joel Hron still positioned CoCounsel as the trusted professional system. MLQ reports that Harvey and Legora are API customers. Your legal vendor may end up running on Astra for Law, the same way CoCounsel's new model runs on open weights it did not train.
What Does the Trusted Access Deal Commit You To?
It commits you to a 30-page agreement with no published price and no published API date. Legal Technology reports a 30-page agreement governing Trusted Access. MLQ confirms there is no published launch date or pricing for the API. LawSites reports that Trusted Access, aimed at Am Law 200 firms, features zero data retention on the API, while ChatGPT AI Hub words it as "may include." Confirm in writing that ZDR applies to your own workspace.
The named early adopters, per Artificial Lawyer, are Latham & Watkins, Ropes & Gray, Cooley and Sullivan & Cromwell. Those firms have the KM staff to run their own evaluations. If your team doesn't, a 54% pass rate means something specific: as ChatGPT AI Hub put it, nearly half of the validation questions did not get an all-pass answer. Every output still needs a lawyer's review. The open question is how much review.
Anthropic took a different route with Claude for Legal on May 12. It shipped twelve practice-area plugins and connectors to vendors including Thomson Reuters, LexisNexis and iManage, plus partner plugins from Harvey and Legora, but no proprietary index. Those are two different bets. OpenAI is betting that owning retrieval is the moat. Anthropic is betting that the incumbents' retrieval, reached through connectors, is good enough.
What Should Legal Ops Do Before Renewal?
Run a matched bake-off on your own matters, and score citations separately from answers. A matched bake-off is a comparison where every candidate gets the same questions, the same reasoning budget and the same grader, so the only thing that varies is the system. Harness choices can flip a ranking as easily as the model can.
This Week:
- Ask your OpenAI account team for the Trusted Access agreement and the ZDR terms for your workspace. Get pricing on paper before any pilot starts.
- Ask your current research vendor, whether CoCounsel, Harvey or Legora, which model and index sit behind its answers today, and whether it plans to route to
gpt-6-astra-law.
This Month:
- Pull 50-100 closed research questions from your own matters where the right answer and the controlling authorities are already known.
- Run them through Astra for Law, your incumbent tool, and one general frontier model given CourtListener access. Grade pass/fail and score completion, not averages.
- Score citations as a separate column: real, controlling and correctly characterized. VLAIR says that is where the systems actually differ.
Before Renewal:
- If Astra for Law does not beat your incumbent by more than the confidence interval on your own set, don't switch. Use the result as pricing leverage.
- If it does win, negotiate for the index, not the seat. The retrieval layer is the part OpenAI built, and the part your vendors may soon resell to you.
The Bottom Line
Vertical AI wrappers keep repeating the same result: the wrapper closes its own model's gap, and the frontier has already moved. Legal research looks like enterprise search did a decade ago. The index was the product, until the index became an API everyone could call. A 15-point lift from better retrieval is worth paying for, as long as you know which comparison it came from.
Astra for Law earned its 54%. Make it earn your budget on your own questions.
Continue Reading
- CoCounsel's New Model Runs on Qwen. Go Read the Card.
- Harvey Swapped Models After Agents Sank Its Margin to -50%
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- Your Bake-Off Ranked Harnesses, Not Models. Re-Run It.
- Agents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.
- RAG Build vs Buy: Buy the Index. Build the Eval Set.
