Arthur
by Arthur AI
Open-source-core observability, evals, guardrails and governance for production AI agents and ML models.
Arthur is an AI observability, evaluation and governance platform that monitors ML models, LLM applications and AI agents in production. It is built for enterprise AI, platform and security teams that need guardrails against hallucination, PII leakage and prompt injection, plus continuous evals and an inventory of every agent running across the company.
Arthur is an AI observability, evaluation and governance platform from Arthur AI, a New York company founded in 2019 and led by CEO and co-founder Adam Wenchel, which has raised more than $60M, including a $15M Series A led by Index Ventures in December 2020 and a $42M Series B co-led by Acrew Capital and Greycroft in September 2022. It began as a machine-learning monitoring tool that tracked accuracy drift, bias and explainability for tabular, NLP and computer-vision models, and has since been rebuilt around generative AI and agents. Its core, the Arthur Engine, was open-sourced under the MIT licence in April 2025 and combines a GenAI Engine (real-time guardrails for PII, hallucination, prompt injection and toxic language, with hallucination checked claim by claim against the supplied context), an ML Engine for drift and accuracy metrics, OpenTelemetry/OpenInference trace collection, prompt management, and continuous evaluations that score groundedness, tool selection, relevance and PII leakage. The commercial Arthur Platform adds dashboards, custom alerting and webhooks, experiments that replay gold-standard datasets against new agent versions, discovery of unregistered shadow agents across endpoints, cloud and private data centres, behavioural runtime security, and token-cost tracking with rerouting to cheaper models, streaming findings to Splunk, Datadog, Elastic and CrowdStrike Falcon. It runs as SaaS, in a dedicated or managed VPC, or self-hosted via Docker Compose. Named users include Axios, The Philadelphia Inquirer, Expel and Upsolve, which used Arthur evals to catch a GPT-5 regression before customers saw it. Its earlier open-source LLM evaluation tool, Arthur Bench, was archived in September 2026. Arthur competes with Arize AI, Fiddler AI, Galileo and LangSmith.
An enterprise AI platform or ML engineering lead in a regulated industry who must ship LLM agents to production and prove they are monitored, evaluated and guarded, ideally without sending data to a third-party SaaS.
Regressions, hallucinations, PII leaks and prompt injections in production agents are caught by continuous evals and real-time guardrails before customers are affected.
At a Glance
- Category
- Governance & Security
- Pricing
- Freemium, Subscription, Contact for pricing
- Target Market
- CIOs, CTOs, CISOs, ML Engineers, Enterprise Developers, Data Scientists
- Deployment
- Cloud-first, Self-hosted, Open-source
- Founded
- 2019
- Headquarters
- New York, USA
Key Features
- ✓Real-time LLM guardrails
Detects and blocks PII, hallucinations, prompt injection and toxic language in model inputs and outputs, reducing the risk of harmful or non-compliant responses reaching users.
- ✓Claim-level hallucination detection
Checks groundedness claim by claim against message history or retrieved context, so RAG and agent answers unsupported by source data are flagged.
- ✓Continuous evals and experiments
Scores agents on groundedness, tool selection, relevance and PII leakage, and replays gold-standard datasets against new versions to surface regressions before release.
- ✓OpenTelemetry agent tracing
Collects OpenInference/OpenTelemetry traces with token counts, latencies and retrieval performance, giving engineers step-level visibility into multi-step agent behaviour.
- ✓Agent discovery and governance
Finds registered and shadow agents running on endpoints, cloud and private data centres and streams findings to SOC tools such as Splunk and CrowdStrike Falcon.
- ✓ML model monitoring
Tracks drift, accuracy metrics and model comparisons for traditional machine-learning models, so classic and generative AI are monitored in one place.
- ✓Cost control
Tracks token spend per application and can reroute traffic to cheaper models, helping teams keep agent operating costs under control as usage scales.
Capabilities
Use Cases
- •Catching model-upgrade regressions
A product team replays gold-standard datasets through its agent after switching LLMs; Upsolve used this to catch a GPT-5 regression before customers noticed.
- •Guarding customer-facing chatbots
An enterprise places Arthur guardrails in front of a support assistant to block prompt injection and redact PII in real time before responses are returned.
- •Governing shadow AI agents
A security team discovers unregistered agents running across endpoints and cloud accounts and routes the findings into its existing SIEM for investigation and policy enforcement.
- •Monitoring production ML model drift
A financial-services data science team monitors credit or fraud models for drift and accuracy decay and receives alerts before business metrics degrade.
- •Self-hosted RAG evaluation
A regulated company deploys the open-source Arthur Engine with Docker inside its own environment to evaluate RAG pipeline groundedness without sending data to SaaS.
Ideal For
Best For
- ✓Adding real-time guardrails (PII, hallucination, prompt injection, toxicity) in front of production LLM applications
- ✓Continuous evaluation and regression testing of AI agents when swapping or upgrading the underlying model
- ✓Self-hosting an open-source AI monitoring engine inside a customer-managed environment for data-governance reasons
- ✓Monitoring classic ML models for drift and accuracy alongside GenAI workloads in one platform
- ✓Inventorying shadow and unregistered AI agents across endpoints and cloud for security and compliance teams
Not Ideal For
- ✗Teams that want a large open-source community and ecosystem: the arthur-engine repository had about 90 GitHub stars when checked in September 2026, far smaller than rival open-source LLM observability projects
- ✗Organisations needing SSO, SLAs, a BAA or data retention beyond 30 days on a small budget, since those sit only in the custom-priced Enterprise tier
- ✗Anyone planning to build on Arthur Bench, the company's earlier open-source LLM evaluation tool, which was archived and made read-only in September 2026
Integrations
Deployment
Market Analysis
Pros
- ✓Open-source MIT-licensed engine that the vendor says sustains sub-second p90 latency at 100+ RPS and has processed 10B+ tokens a month for Fortune 100 customers
- ✓Broad coverage from classic ML drift monitoring through LLM guardrails, agent tracing, evals and agent discovery
- ✓Transparent published pricing with a usable free tier and unlimited seats
- ✓SOC 2 Type II audited (2022) and deployable in a customer VPC for regulated industries
Cons
- ✗Small open-source footprint: about 90 GitHub stars on arthur-engine (September 2026) and a modest Hacker News launch (22 points), so community support is thin
- ✗Arthur Bench, its earlier open-source LLM evaluation tool, was archived in September 2026, a signal of product churn for teams that adopted it
- ✗Free and Premium tiers cap retention at 7 and 30 days and limit spans, inferences and evals; SSO and BAA are Enterprise-only
- ✗Little independent review data is publicly readable; the G2 page blocked access and no Capterra listing was found
Pricing
Free
$0
- ✓Up to 4 use cases
- ✓Unlimited seats
- ✓7-day data retention
- ✓5k jobs, 300k spans, 12k inferences, 3k evals
- ✓3 custom alerts per model
- ✓API and UI access
Premium
From $60/mo
- ✓Up to 100 use cases
- ✓30-day data retention
- ✓20k jobs, 1.2M spans, 100k inferences, 75k evals
- ✓Customizable metrics and dashboards
- ✓Unlimited custom alerts and webhook integrations
- ✓10 projects
Enterprise
Contact for pricing
- ✓Dedicated or managed VPC options
- ✓Unlimited data retention and custom limits
- ✓SSO, SLAs and BAA
- ✓Dedicated customer success manager
- ✓Unlimited organizations, workspaces and projects
List pricing is published: a Free tier (4 use cases, 7-day retention) and Premium at $60/month, metered by use cases, jobs, spans, inferences and evals rather than seats (seats are unlimited). SSO, SLAs, a BAA, VPC deployment and unlimited retention require custom-priced Enterprise. The Arthur Engine itself is free and open source under the MIT licence.
Security & Compliance
Connect
Sources
This page was written from 10 sources, 6 on domains other than arthur.ai.
- 1.arthur.ai — arthur.aivendor
- 2.arthur.ai — pricingvendor
- 3.github.com — arthur engine
- 4.github.com — bench
- 5.techcrunch.com — arthur ai snags 15m series a to grow machine learning monito
- 6.prnewswire.com — arthur raises 42m in series b funding as ai adoption soars 3
- 7.news.ycombinator.com — item
- 8.arthur.ai — how upsolve built trusted agentic ai with arthurvendor
- 9.arthur.ai — arthur achieves soc 2 r type ii certification compliancevendor
- 10.docs.arthur.ai — docs.arthur.ai
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Baselayer
Business identity, risk and fraud infrastructure that now verifies AI agents before they transact
SAS AI Navigator
SaaS AI governance inventory for models, agents and use cases, sold through Microsoft Marketplace
Rein Security
Runtime security for enterprise AI agents, deployed as a sidecar inside your own cloud
Armadin
Autonomous offensive security: AI agent swarms that chain your weak spots into proven attack paths before an attacker does