Vals AI
by Vals
Independent AI benchmarking on private test sets, plus custom evals built from your own code, for legal, finance, healthcare and coding workloads
Vals AI is an independent AI evaluation company that benchmarks frontier models and AI applications on economically valuable tasks in law, finance, healthcare, coding and public services, using private test sets that vendors cannot train against. It is for AI leaders, legal and finance teams and model builders who need neutral evidence before choosing or shipping a model.
Vals AI is a San Francisco evaluation company founded in 2024 and led by co-founder and CEO Rayan Krishnan that measures how frontier models and AI applications perform on economically valuable work rather than academic trivia. Its core design choice is secrecy: Vals runs every evaluation in-house against private, domain-expert test sets in law, finance, tax, healthcare, coding and public services, so model makers cannot train against the answers. The public face of the company is a set of free leaderboards, including the Vals Index (a composite across coding, finance and legal tasks), a Multimodal Index, and benchmarks such as Finance Agent v2, CorpFin v2, CaseLaw v2, LegalBench, Legal Research Bench, MedCode, MedScribe, TaxEval, ProofBench, Terminal-Bench and SWE-bench Verified. In February 2025 it published the Vals Legal AI Report with law firms including Reed Smith and Fisher Phillips, followed in October 2025 by a legal research study in which Counsel Stack (81%), Alexi and ChatGPT (80%) and Midpage (79%) all beat a 71% lawyer baseline. Commercially, Vals sells private Test Suites with custom checks and CI/CD integration, and Vals Smith, which turns merged pull requests from a customer's GitHub repository into hidden-test tasks so teams can compare models and coding agents on their own code; Vercel, Exa, Fireworks AI and Harvey are listed as users. A public-sector program includes a SNAP benefits benchmark built with Code for America. In September 2026 Vals published evidence of rising benchmark cheating by frontier models, and TechCrunch reported a $40 million Series A led by Andreessen Horowitz, following a seed round led by 8VC and Bloomberg Beta, with revenue up eightfold year on year and headcount growing from 8 to 25.
Heads of AI or AI governance leads at regulated enterprises (legal, finance, healthcare) who must justify which model or vertical AI vendor they deploy with neutral, domain-specific evidence.
Model and vendor decisions backed by independent scores on private, domain-specific tasks, including on your own codebase, instead of vendor-reported benchmarks.
At a Glance
- Category
- Governance & Security
- Pricing
- Contact for pricing, Subscription, Usage-based
- Target Market
- CTOs, Heads of AI, AI Governance Teams, Legal Operations, Model Builders
- Deployment
- Cloud-first, API-based
- Founded
- 2024
- Headquarters
- San Francisco, United States
- Team Size
- 11-50
Key Features
- ✓Vals Index
A composite benchmark across coding, finance and legal tasks that gives buyers one comparable score for frontier models.
- ✓Private domain benchmarks
Held-out test sets in law, finance, tax, healthcare and coding that model vendors cannot train against, reducing contamination.
- ✓Vals Smith
Converts merged pull requests from a public or private GitHub repository into hidden-test tasks to benchmark models and coding agents on your own code.
- ✓Enterprise Test Suites
Private test suites with custom checks, hallucination detection and CI/CD integration for regression testing across model upgrades.
- ✓Vertical AI industry reports
Blind, rubric-graded head-to-head studies of legal AI vendors run with law firms and published openly for buyers.
- ✓Benchmark-integrity audits
Independent analysis of models attempting to cheat on benchmarks, exposing gaps between vendor-claimed and independently measured scores.
Capabilities
Use Cases
- •Selecting a legal research tool
A law firm compares vendors using Vals' published legal research study, where Counsel Stack scored 81% against a 71% lawyer baseline.
- •Choosing a coding agent on your own code
An engineering organization connects its GitHub repository to Vals Smith and compares pass rates of models and agents on real historical fixes.
- •Gating model upgrades
A platform team runs private test suites in CI before switching models, catching regressions on domain tasks before they reach production.
- •Verifying vendor benchmark claims
An AI governance lead checks independent scores, such as Vals measuring Gemini 3.8 Flash at 71.7% against a claimed 88.8%, before approval.
- •Public-sector AI procurement
A benefits agency uses the SNAP Public Benefits Benchmark to judge whether any model is accurate enough for citizen-facing guidance.
Ideal For
Best For
- ✓Legal and finance teams choosing between vertical AI vendors using neutral, published head-to-head results
- ✓Engineering leaders deciding which coding model or agent to standardize on, using Vals Smith benchmarks built from their own merged pull requests
- ✓Model builders and AI vendors that need a third-party score buyers will trust more than self-reported benchmarks
- ✓Government agencies assessing AI for public-facing services such as benefits guidance
- ✓AI governance teams that want evidence of regressions or benchmark gaming before approving a model upgrade
Not Ideal For
- ✗Teams that want a self-serve eval tool with published pricing today; Vals sells through demos and negotiated contracts
- ✗Organizations that require fully reproducible, open test sets for audit, since Vals keeps its test materials private by design
- ✗Teams whose main need is production LLM tracing and observability; full-stack platforms such as LangSmith, Langfuse or Braintrust are closer fits
Integrations
Deployment
Market Analysis
Pros
- ✓Independent third-party evaluation with private test sets, which resists the benchmark contamination that undermines public leaderboards
- ✓Rare depth in regulated domains: legal, finance, tax and healthcare benchmarks built with practitioners
- ✓Vals Smith lets teams benchmark models and agents on their own code rather than generic tasks
- ✓Strong momentum: a16z-led $40M Series A, revenue up 8x year on year and published cheating audits that drew wide attention
Cons
- ✗Potential conflict of interest: AI labs pay Vals to be evaluated, and the commercial model sells to the same companies it ranks
- ✗Private test sets mean results cannot be independently reproduced, and Hacker News commenters have questioned whether rankings match observed utility and whether every score was run in-house
- ✗Vendor opt-outs limit comparisons: Thomson Reuters and LexisNexis declined its 2025 legal research study and vLex withdrew before publication
- ✗No public pricing, no published security certifications, and Vals Smith currently covers coding on GitHub only
Pricing
Public benchmarks and leaderboards
$0
- ✓Vals Index and Multimodal Index
- ✓Legal, finance, healthcare, tax and coding benchmarks
- ✓Published industry reports
Enterprise (Test Suites, Vals Smith, custom evaluations)
Contact for pricing
- ✓Private domain test suites with custom checks
- ✓Vals Smith custom coding benchmarks from GitHub repositories
- ✓CI/CD integration and SDK/CLI
- ✓Expert review workflows
Vals publishes no list pricing. Its public leaderboards and industry reports are free to read, while enterprise access to Test Suites, Vals Smith and custom domain evaluations is sold through demo-led, negotiated contracts combining a platform subscription with usage-based evaluation volume, according to Sacra. Separately, AI labs pay Vals to have their models evaluated, which CEO Rayan Krishnan likens to paying the College Board to sit the SAT.
Security & Compliance
Connect
Sources
This page was written from 11 sources, 7 on domains other than vals.ai.
- 1.vals.ai — vals.aivendor
- 2.vals.ai — vals smithvendor
- 3.vals.ai — govvendor
- 4.vals.ai — cheating on the risevendor
- 5.techcrunch.com — vals backed by andreessen horowitz is looking to become the
- 6.lawnext.com — vals ais latest benchmark finds legal and general ai now out
- 7.legaltechnologyhub.com — vals ai
- 8.sacra.com — vals ai
- 9.github.com — vals ai
- 10.everydev.ai — vals ai
- 11.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Alation AIOS
Governed data, context and AI agents in one intelligence operating system for the enterprise
Apollo Research Watcher
Runtime monitoring and blocking for Claude Code and Codex: catch dangerous coding-agent actions before they run
Comp AI
Open-source, agentic compliance automation for SOC 2, ISO 27001, HIPAA and GDPR: an AGPL alternative to Vanta and Drata
Mate Security
Open agentic SOC platform powered by a security context graph built for each organisation