Patronus AI
by Patronus AI
Evaluation, agent debugging and simulation environments for reliable AI agents
Patronus AI sells an evaluation platform for LLM and agent applications — managed evaluators, the Percival trace-debugging agent and open judge models such as Lynx and GLIDER — and, more recently, RL environments and Digital World Models for training and testing AI agents on simulated digital workflows. It targets AI engineering teams and frontier labs.
Patronus AI was founded in 2023 by Anand Kannappan (CEO), who helped develop explainable ML frameworks at Meta Reality Labs, and Rebecca Qian, who led responsible NLP research at Meta AI, and launched with a $3 million seed round led by Lightspeed Venture Partners to evaluate and test LLMs for regulated industries. Its evaluation platform remains on sale: a Python SDK and web app for experiments, logs and traces, plus a managed evaluator library — Glider for fast guardrail and rubric checks, Judge for customizable LLM-as-a-judge scoring, Judge MM for image and audio, and Lynx for hallucinations — covering context relevance, context sufficiency, answer relevance, enterprise PII and toxicity. The company publishes much of its research openly: Lynx, a Llama-3 fine-tune released in 8B and 70B sizes in July 2024, whose 70B version scored 87.4% on the company's HaluBench versus 86.5% for GPT-4o, and GLIDER, a 4B Phi-3.5-mini judge released in December 2024 that grades against user-defined rubrics; the Lynx 70B and GLIDER weights carry a non-commercial CC-BY-NC-4.0 licence. Percival, an agent that reads traces from LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI and smolagents, detects 20+ failure modes and suggests fixes. Since late 2025 the emphasis has shifted toward training-side simulation: Generative Simulators, RL environments that co-generate tasks, world dynamics and reward functions, arrived in December 2025, followed by Digital World Models — language diffusion world models for agent training and evaluation — announced with a $50 million Series B in June 2026 led by Greenfield Partners, bringing total funding to $70 million. Patronus says revenue grew more than 15x over the prior year and that it works with the majority of leading frontier AI labs and hyperscalers, placing it against evaluation vendors such as Arize, Galileo and Braintrust as well as agent-training data suppliers.
AI platform or ML engineering leads at regulated enterprises shipping RAG and agent applications who want managed hallucination and agent-failure evaluators instead of maintaining their own LLM-as-a-judge infrastructure.
Automated, research-backed scoring of hallucinations and agent failures before and after release, so unreliable outputs are caught before customers see them.
At a Glance
- Category
- Developer Tools
- Pricing
- Freemium, Usage-based, Subscription, Contact for pricing
- Target Market
- CTOs, AI/ML Engineers, Enterprise Developers, AI Research Labs
- Deployment
- Cloud-first, API-based
- Founded
- 2023
- Headquarters
- San Francisco, USA
Key Features
- ✓Patronus Evaluators
Managed evaluator library — Glider, Judge, multimodal Judge MM and Lynx — scoring hallucination, context relevance, answer relevance, PII and toxicity without building judge infrastructure.
- ✓Lynx hallucination model
Open-weight Llama-3 fine-tune in 8B and 70B sizes that returns PASS or FAIL with reasoning on whether an answer is faithful to its source documents.
- ✓GLIDER judge model
4B-parameter Phi-3.5-mini fine-tune that grades text, conversations and RAG output against user-defined rubrics, returning reasoning, highlighted key phrases and an integer score.
- ✓Percival agent debugger
Agent that analyses traces from LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI and smolagents, detects 20+ failure modes and suggests fixes, improving as users annotate issues.
- ✓Experiments, logs and tracing
MIT-licensed Python SDK and web platform for tracing functions, running experiments and comparing LLM application versions so quality regressions surface before release.
- ✓RL environments (Generative Simulators)
Adaptive environments that co-generate tasks, world dynamics and reward functions for training agents on long-horizon, real-world-like digital workflows.
- ✓Digital World Models
Language diffusion world models, unveiled June 2026, that scale creation of simulation data to train and evaluate agent actions across complex digital workflows.
Capabilities
Use Cases
- •RAG faithfulness gating
A financial-services RAG assistant runs Lynx or hallucination evaluators on each answer, flagging responses unsupported by the retrieved documents before analysts rely on them.
- •Agent failure debugging
An engineering team sends LangGraph or crewAI traces to Percival, which identifies failure modes across 20+ categories and proposes optimizations to the agent workflow.
- •Release regression testing
Before shipping a new model or prompt, teams run experiments against curated datasets and compare evaluator scores with the previous version to catch quality regressions.
- •Custom policy and PII checks
Compliance and platform teams define rubric-based Glider or Judge evaluators for company policy, style and enterprise PII, then use them as guardrails on production outputs.
- •Frontier agent training
A frontier lab trains long-horizon agents on Patronus RL environments and Digital World Models that simulate software, research, communication and enterprise workflows.
Ideal For
Best For
- ✓Hallucination and faithfulness scoring for RAG applications in finance, healthcare or legal settings
- ✓Debugging multi-step agent failures from LangGraph, crewAI or OpenAI Agents SDK traces with Percival
- ✓Running pre-release LLM experiments and version-to-version regression comparisons with managed evaluators
- ✓Custom rubric-based judges for policy, tone and PII checks without building judge infrastructure
- ✓Frontier AI labs sourcing RL environments and simulation data for long-horizon agent training
Not Ideal For
- ✗Companies that want to self-host open-weight evaluators commercially, because the Lynx 70B and GLIDER weights are licensed CC-BY-NC-4.0 (non-commercial).
- ✗Teams wanting long history on a free plan: the Developer tier keeps only the last two weeks of experiments, logs and traces and caps usage at two projects.
- ✗Enterprise app teams expecting self-serve access to RL environments or Digital World Models, which have no published pricing and are sold by contacting the company.
Integrations
Deployment
Market Analysis
Pros
- ✓Published, usage-based pricing and a free Developer tier make it cheap to pilot managed evaluators before a procurement cycle.
- ✓Research-backed evaluators: the Lynx 70B model scored 87.4% on HaluBench versus 86.5% for GPT-4o in the company's own published evaluation.
- ✓Percival works across the major agent frameworks (LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI, smolagents) and flags 20+ failure modes from traces.
- ✓Enterprise tier offers on-prem or dedicated VPC deployment with SSO and custom data retention, which suits regulated industries.
- ✓Well capitalised, with $70M raised including a $50M Series B in June 2026 and reported revenue growth of more than 15x year over year.
Cons
- ✗The company's 2025-2026 launches (Generative Simulators, Digital World Models, post-training datasets) target frontier-lab agent training rather than enterprise app evaluation, a roadmap-risk signal for evaluation-only buyers.
- ✗The open Lynx 70B and GLIDER weights are CC-BY-NC-4.0 non-commercial, so enterprises cannot self-host them in production without a separate arrangement.
- ✗Little independent practitioner feedback exists: Hacker News threads about the company are sparse and low-engagement, and no verifiable review-site rating was found.
- ✗The Python SDK was still pre-1.0 (version 0.1.25, January 2026 on PyPI), and simulation products have no public pricing or self-serve access.
Pricing
Developer
$0
- ✓No credit card required
- ✓2 projects, 5 experiments per project
- ✓Last 2 weeks of experiments, logs and traces
- ✓Unlimited comparisons and datasets
- ✓Optional evaluator API: $10 / 1k small evaluator calls, $20 / 1k large evaluator calls, $10 / 1k eval explanations
Base
From $25/mo
- ✓600 pages included
- ✓Page add-ons on demand
Enterprise
Contact for pricing
- ✓Everything unlimited
- ✓On-prem or dedicated VPC deployment
- ✓Custom data retention
- ✓SSO
- ✓Custom eval model fine-tuning and eval dataset generation
The evaluation platform publishes list pricing: a free Developer tier with no credit card, a $25/month Base plan, and metered evaluator API calls at $10 per 1,000 small-evaluator calls, $20 per 1,000 large-evaluator calls and $10 per 1,000 eval explanations. On-prem or dedicated VPC deployment, SSO, custom data retention and custom evaluator fine-tuning are Enterprise-only and quoted on request. RL environments and Digital World Models have no published pricing.
Security & Compliance
Connect
Sources
This page was written from 15 sources, 9 on domains other than patronus.ai.
- 1.patronus.ai — patronus.aivendor
- 2.patronus.ai — pricingvendor
- 3.patronus.ai — percivalvendor
- 4.patronus.ai — rl environmentsvendor
- 5.patronus.ai — blogvendor
- 6.patronus.ai — patronus evaluatorsvendor
- 7.prnewswire.com — patronus ai raises 50 million series b and unveils first dig
- 8.techcrunch.com — patronus ai conjures up an llm evaluation tool for regulated
- 9.venturebeat.com — meet patronus ais lynx the open source bullshit detector out
- 10.huggingface.co — Llama 3 Patronus Lynx 70B Instruct
- 11.huggingface.co — glider
- 12.huggingface.co — PatronusAI
- 13.github.com — patronus ai
- 14.pypi.org — patronus
- 15.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
HackerRank Chakra
AI interviewer that runs hands-on technical interviews and scores how candidates work with AI
Autoheal
Self-improving AI agents for incident response, vulnerability fixes and release work after code is written
Unified.to
Real-time unified API and MCP server giving AI products and agents access to 1,000+ SaaS integrations
Momentic
Agentic QA platform that writes, runs and self-heals end-to-end tests for web and mobile apps