P

Patronus AI

by Patronus AI

Developer ToolsAgent DevelopmentGovernance & SecurityAI Models & APIs

Evaluation, agent debugging and simulation environments for reliable AI agents

Freemium · Usage-based · Subscription · Contact for pricing·Added Jul 10, 2026·Updated Sep 14, 2026
Share:
THE DAILY BRIEF
Patronus AI

by Patronus AI

Developer ToolsAgent DevelopmentGovernance & SecurityAI Models & APIs

Evaluation, agent debugging and simulation environments for reliable AI agents

Freemium · Usage-based · Subscription · Contact for pricing

Patronus AI sells an evaluation platform for LLM and agent applications — managed evaluators, the Percival trace-debugging agent and open judge models such as Lynx and GLIDER — and, more recently, RL environments and Digital World Models for training and testing AI agents on simulated digital workflows. It targets AI engineering teams and frontier labs.

At a Glance

Category
Developer Tools
Pricing
Freemium, Usage-based, Subscription, Contact for pricing
Target Market
CTOs, AI/ML Engineers, Enterprise Developers, AI Research Labs
Deployment
Cloud-first, API-based
Founded
2023
Headquarters
San Francisco, USA

Key Features

  • ✓Patronus Evaluators
  • ✓Lynx hallucination model
  • ✓GLIDER judge model
  • ✓Percival agent debugger
  • ✓Experiments, logs and tracing
  • ✓RL environments (Generative Simulators)
  • ✓Digital World Models

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✓fine tuning
✗agent orchestration

Use Cases

  • •RAG faithfulness gating
  • •Agent failure debugging
  • •Release regression testing
  • •Custom policy and PII checks
  • •Frontier agent training

Ideal For

Best For

  • ✓Hallucination and faithfulness scoring for RAG applications in finance, healthcare or legal settings
  • ✓Debugging multi-step agent failures from LangGraph, crewAI or OpenAI Agents SDK traces with Percival
  • ✓Running pre-release LLM experiments and version-to-version regression comparisons with managed evaluators
  • ✓Custom rubric-based judges for policy, tone and PII checks without building judge infrastructure
  • ✓Frontier AI labs sourcing RL environments and simulation data for long-horizon agent training

Not Ideal For

  • ✗Companies that want to self-host open-weight evaluators commercially, because the Lynx 70B and GLIDER weights are licensed CC-BY-NC-4.0 (non-commercial).
  • ✗Teams wanting long history on a free plan: the Developer tier keeps only the last two weeks of experiments, logs and traces and caps usage at two projects.
  • ✗Enterprise app teams expecting self-serve access to RL environments or Digital World Models, which have no published pricing and are sold by contacting the company.

Market Analysis

AI evaluation and reliabilityResearch-ledFrontier-lab agent training infrastructure

Pros

  • ✓Published, usage-based pricing and a free Developer tier make it cheap to pilot managed evaluators before a procurement cycle.
  • ✓Research-backed evaluators: the Lynx 70B model scored 87.4% on HaluBench versus 86.5% for GPT-4o in the company's own published evaluation.
  • ✓Percival works across the major agent frameworks (LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI, smolagents) and flags 20+ failure modes from traces.
  • ✓Enterprise tier offers on-prem or dedicated VPC deployment with SSO and custom data retention, which suits regulated industries.
  • ✓Well capitalised, with $70M raised including a $50M Series B in June 2026 and reported revenue growth of more than 15x year over year.

Cons

  • ✗The company's 2025-2026 launches (Generative Simulators, Digital World Models, post-training datasets) target frontier-lab agent training rather than enterprise app evaluation, a roadmap-risk signal for evaluation-only buyers.
  • ✗The open Lynx 70B and GLIDER weights are CC-BY-NC-4.0 non-commercial, so enterprises cannot self-host them in production without a separate arrangement.
  • ✗Little independent practitioner feedback exists: Hacker News threads about the company are sparse and low-engagement, and no verifiable review-site rating was found.
  • ✗The Python SDK was still pre-1.0 (version 0.1.25, January 2026 on PyPI), and simulation products have no public pricing or self-serve access.

Pricing

Developer

$0

  • ✓No credit card required
  • ✓2 projects, 5 experiments per project
  • ✓Last 2 weeks of experiments, logs and traces
  • ✓Unlimited comparisons and datasets
  • ✓Optional evaluator API: $10 / 1k small evaluator calls, $20 / 1k large evaluator calls, $10 / 1k eval explanations

Base

From $25/mo

  • ✓600 pages included
  • ✓Page add-ons on demand

Enterprise

Contact for pricing

  • ✓Everything unlimited
  • ✓On-prem or dedicated VPC deployment
  • ✓Custom data retention
  • ✓SSO
  • ✓Custom eval model fine-tuning and eval dataset generation

The evaluation platform publishes list pricing: a free Developer tier with no credit card, a $25/month Base plan, and metered evaluator API calls at $10 per 1,000 small-evaluator calls, $20 per 1,000 large-evaluator calls and $10 per 1,000 eval explanations. On-prem or dedicated VPC deployment, SSO, custom data retention and custom evaluator fine-tuning are Enterprise-only and quoted on request. RL environments and Digital World Models have no published pricing.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✓sso
✗data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Patronus AI sells an evaluation platform for LLM and agent applications — managed evaluators, the Percival trace-debugging agent and open judge models such as Lynx and GLIDER — and, more recently, RL environments and Digital World Models for training and testing AI agents on simulated digital workflows. It targets AI engineering teams and frontier labs.

Patronus AI was founded in 2023 by Anand Kannappan (CEO), who helped develop explainable ML frameworks at Meta Reality Labs, and Rebecca Qian, who led responsible NLP research at Meta AI, and launched with a $3 million seed round led by Lightspeed Venture Partners to evaluate and test LLMs for regulated industries. Its evaluation platform remains on sale: a Python SDK and web app for experiments, logs and traces, plus a managed evaluator library — Glider for fast guardrail and rubric checks, Judge for customizable LLM-as-a-judge scoring, Judge MM for image and audio, and Lynx for hallucinations — covering context relevance, context sufficiency, answer relevance, enterprise PII and toxicity. The company publishes much of its research openly: Lynx, a Llama-3 fine-tune released in 8B and 70B sizes in July 2024, whose 70B version scored 87.4% on the company's HaluBench versus 86.5% for GPT-4o, and GLIDER, a 4B Phi-3.5-mini judge released in December 2024 that grades against user-defined rubrics; the Lynx 70B and GLIDER weights carry a non-commercial CC-BY-NC-4.0 licence. Percival, an agent that reads traces from LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI and smolagents, detects 20+ failure modes and suggests fixes. Since late 2025 the emphasis has shifted toward training-side simulation: Generative Simulators, RL environments that co-generate tasks, world dynamics and reward functions, arrived in December 2025, followed by Digital World Models — language diffusion world models for agent training and evaluation — announced with a $50 million Series B in June 2026 led by Greenfield Partners, bringing total funding to $70 million. Patronus says revenue grew more than 15x over the prior year and that it works with the majority of leading frontier AI labs and hyperscalers, placing it against evaluation vendors such as Arize, Galileo and Braintrust as well as agent-training data suppliers.

Ideal Buyer

AI platform or ML engineering leads at regulated enterprises shipping RAG and agent applications who want managed hallucination and agent-failure evaluators instead of maintaining their own LLM-as-a-judge infrastructure.

Key Benefit

Automated, research-backed scoring of hallucinations and agent failures before and after release, so unreliable outputs are caught before customers see them.

At a Glance

Category
Developer Tools
Pricing
Freemium, Usage-based, Subscription, Contact for pricing
Target Market
CTOs, AI/ML Engineers, Enterprise Developers, AI Research Labs
Deployment
Cloud-first, API-based
Founded
2023
Headquarters
San Francisco, USA

Key Features

  • ✓
    Patronus Evaluators

    Managed evaluator library — Glider, Judge, multimodal Judge MM and Lynx — scoring hallucination, context relevance, answer relevance, PII and toxicity without building judge infrastructure.

  • ✓
    Lynx hallucination model

    Open-weight Llama-3 fine-tune in 8B and 70B sizes that returns PASS or FAIL with reasoning on whether an answer is faithful to its source documents.

  • ✓
    GLIDER judge model

    4B-parameter Phi-3.5-mini fine-tune that grades text, conversations and RAG output against user-defined rubrics, returning reasoning, highlighted key phrases and an integer score.

  • ✓
    Percival agent debugger

    Agent that analyses traces from LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI and smolagents, detects 20+ failure modes and suggests fixes, improving as users annotate issues.

  • ✓
    Experiments, logs and tracing

    MIT-licensed Python SDK and web platform for tracing functions, running experiments and comparing LLM application versions so quality regressions surface before release.

  • ✓
    RL environments (Generative Simulators)

    Adaptive environments that co-generate tasks, world dynamics and reward functions for training agents on long-horizon, real-world-like digital workflows.

  • ✓
    Digital World Models

    Language diffusion world models, unveiled June 2026, that scale creation of simulation data to train and evaluate agent actions across complex digital workflows.

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✓fine tuning
✗agent orchestration

Use Cases

  • •
    RAG faithfulness gating

    A financial-services RAG assistant runs Lynx or hallucination evaluators on each answer, flagging responses unsupported by the retrieved documents before analysts rely on them.

  • •
    Agent failure debugging

    An engineering team sends LangGraph or crewAI traces to Percival, which identifies failure modes across 20+ categories and proposes optimizations to the agent workflow.

  • •
    Release regression testing

    Before shipping a new model or prompt, teams run experiments against curated datasets and compare evaluator scores with the previous version to catch quality regressions.

  • •
    Custom policy and PII checks

    Compliance and platform teams define rubric-based Glider or Judge evaluators for company policy, style and enterprise PII, then use them as guardrails on production outputs.

  • •
    Frontier agent training

    A frontier lab trains long-horizon agents on Patronus RL environments and Digital World Models that simulate software, research, communication and enterprise workflows.

Ideal For

Best For

  • ✓Hallucination and faithfulness scoring for RAG applications in finance, healthcare or legal settings
  • ✓Debugging multi-step agent failures from LangGraph, crewAI or OpenAI Agents SDK traces with Percival
  • ✓Running pre-release LLM experiments and version-to-version regression comparisons with managed evaluators
  • ✓Custom rubric-based judges for policy, tone and PII checks without building judge infrastructure
  • ✓Frontier AI labs sourcing RL environments and simulation data for long-horizon agent training

Not Ideal For

  • ✗Companies that want to self-host open-weight evaluators commercially, because the Lynx 70B and GLIDER weights are licensed CC-BY-NC-4.0 (non-commercial).
  • ✗Teams wanting long history on a free plan: the Developer tier keeps only the last two weeks of experiments, logs and traces and caps usage at two projects.
  • ✗Enterprise app teams expecting self-serve access to RL environments or Digital World Models, which have no published pricing and are sold by contacting the company.

Integrations

✓SDK Available
SDK:Python

Deployment

✓On-Premise

Market Analysis

AI evaluation and reliabilityResearch-ledFrontier-lab agent training infrastructure

Pros

  • ✓Published, usage-based pricing and a free Developer tier make it cheap to pilot managed evaluators before a procurement cycle.
  • ✓Research-backed evaluators: the Lynx 70B model scored 87.4% on HaluBench versus 86.5% for GPT-4o in the company's own published evaluation.
  • ✓Percival works across the major agent frameworks (LangGraph, crewAI, OpenAI Agents SDK, Pydantic AI, smolagents) and flags 20+ failure modes from traces.
  • ✓Enterprise tier offers on-prem or dedicated VPC deployment with SSO and custom data retention, which suits regulated industries.
  • ✓Well capitalised, with $70M raised including a $50M Series B in June 2026 and reported revenue growth of more than 15x year over year.

Cons

  • ✗The company's 2025-2026 launches (Generative Simulators, Digital World Models, post-training datasets) target frontier-lab agent training rather than enterprise app evaluation, a roadmap-risk signal for evaluation-only buyers.
  • ✗The open Lynx 70B and GLIDER weights are CC-BY-NC-4.0 non-commercial, so enterprises cannot self-host them in production without a separate arrangement.
  • ✗Little independent practitioner feedback exists: Hacker News threads about the company are sparse and low-engagement, and no verifiable review-site rating was found.
  • ✗The Python SDK was still pre-1.0 (version 0.1.25, January 2026 on PyPI), and simulation products have no public pricing or self-serve access.

Pricing

Developer

$0

  • ✓No credit card required
  • ✓2 projects, 5 experiments per project
  • ✓Last 2 weeks of experiments, logs and traces
  • ✓Unlimited comparisons and datasets
  • ✓Optional evaluator API: $10 / 1k small evaluator calls, $20 / 1k large evaluator calls, $10 / 1k eval explanations

Base

From $25/mo

  • ✓600 pages included
  • ✓Page add-ons on demand

Enterprise

Contact for pricing

  • ✓Everything unlimited
  • ✓On-prem or dedicated VPC deployment
  • ✓Custom data retention
  • ✓SSO
  • ✓Custom eval model fine-tuning and eval dataset generation

The evaluation platform publishes list pricing: a free Developer tier with no credit card, a $25/month Base plan, and metered evaluator API calls at $10 per 1,000 small-evaluator calls, $20 per 1,000 large-evaluator calls and $10 per 1,000 eval explanations. On-prem or dedicated VPC deployment, SSO, custom data retention and custom evaluator fine-tuning are Enterprise-only and quoted on request. RL environments and Digital World Models have no published pricing.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✓sso
✗data residency

Connect

Sources

This page was written from 15 sources, 9 on domains other than patronus.ai.

  1. 1.patronus.ai — patronus.aivendor
  2. 2.patronus.ai — pricingvendor
  3. 3.patronus.ai — percivalvendor
  4. 4.patronus.ai — rl environmentsvendor
  5. 5.patronus.ai — blogvendor
  6. 6.patronus.ai — patronus evaluatorsvendor
  7. 7.prnewswire.com — patronus ai raises 50 million series b and unveils first dig
  8. 8.techcrunch.com — patronus ai conjures up an llm evaluation tool for regulated
  9. 9.venturebeat.com — meet patronus ais lynx the open source bullshit detector out
  10. 10.huggingface.co — Llama 3 Patronus Lynx 70B Instruct
  11. 11.huggingface.co — glider
  12. 12.huggingface.co — PatronusAI
  13. 13.github.com — patronus ai
  14. 14.pypi.org — patronus
  15. 15.hn.algolia.com — search
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe