LangWatch
by LangWatch
Open-source agent testing platform pairing simulation-based QA with OpenTelemetry tracing and evals
LangWatch is an Apache-2.0 licensed LLMOps platform for teams putting AI agents into production. It combines OpenTelemetry-native tracing, online and offline evaluation, and end-to-end agent simulation — running scripted user personas against your agent before release — so quality problems are caught in CI rather than by customers.
LangWatch is an Amsterdam-based LLMOps platform founded by Manouk Draisma and Rogerio Chaves that positions itself less as a dashboard and more as a testing loop for AI agents. Its distinguishing capability is simulation: a scenario defines the agent under test, a user simulator that plays a realistic persona across text or voice, and an LLM judge that scores the resulting conversation — so a multi-turn agent can be regression-tested in CI the way deterministic software is, and any observed production failure can be converted into a simulation that verifies the fix and gates the release. Around that sit OpenTelemetry-native tracing over OTLP following the GenAI semantic conventions, dataset creation and manual annotation tools for domain experts, online evaluation scoring live production traffic, automatic prompt optimisation drawing on Stanford's DSPy framework, and an AI gateway for governance and cost control that the project measures at roughly 700 nanoseconds of hot-path overhead. The platform is framework-agnostic, with documented support for LangChain, LangGraph, CrewAI, Google ADK, Mastra, the Vercel AI SDK, LangFlow, Flowise and n8n, and model providers including OpenAI, Anthropic, Azure OpenAI, Vertex AI, Bedrock, Groq and Ollama. It is open-core: the core is Apache 2.0 with 3.5k GitHub stars, the SDKs are MIT-licensed, and enterprise modules require a commercial licence. Self-hosting via Docker Compose, Kubernetes Helm or full on-premises deployment is available on every tier, which is unusual in this category. LangWatch is ISO 27001 certified and GDPR compliant, names Deloitte, Backbase, PagBank, Visma and Freeday as customers, and raised a EUR 1 million pre-seed led by Passion Capital in February 2025.
The AI engineering or QA lead at a regulated European enterprise who needs multi-turn agents regression-tested before release and must keep trace data inside their own infrastructure.
Agent regressions get caught by simulated users in CI before release, instead of being discovered in production traces after customers have already hit them.
At a Glance
- Category
- Agent Development
- Pricing
- Freemium, Subscription, Usage-based, Contact for pricing
- Target Market
- CTOs, AI Engineers, QA and Product Leads, Enterprise Developers
- Deployment
- Open-source, Self-hosted, Cloud-first, Hybrid
- Headquarters
- Amsterdam, Netherlands
Key Features
- ✓Agent simulation testing
Scripted user personas run realistic text or voice scenarios against your agent before it ever reaches production.
- ✓OpenTelemetry-native tracing
Follows GenAI semantic conventions over OTLP, so instrumentation stays portable rather than locked to one vendor.
- ✓Online and offline evaluation
LLM judges, custom code and workflow evaluators score both curated datasets and live production traffic.
- ✓DSPy-based prompt optimisation
Automatically searches for better prompts using Stanford's DSPy approach instead of manual trial and error.
- ✓AI gateway
OpenAI- and Anthropic-compatible proxy adding governance and cost control at roughly 700 nanoseconds hot-path overhead.
- ✓Self-hosting on every tier
Docker Compose, Kubernetes Helm, on-premises and hybrid deployment are available without an enterprise contract.
- ✓Langy AI engineer
Turns product goals into test plans and scenarios, scores them with judge rubrics and opens pull requests with fixes.
Capabilities
Use Cases
- •Pre-release agent QA
Run a suite of simulated customer conversations against a new agent build and block the release if scores drop.
- •Turning incidents into tests
Convert a real production failure trace into a repeatable simulation that proves the fix and stops the regression returning.
- •Regulated deployment in the EU
Self-host the data plane on-premises under ISO 27001 and GDPR while using the managed control plane for the interface.
- •Cost and quality monitoring
Track token spend per conversation alongside evaluator scores to see where quality drops as cheaper models are substituted.
- •Domain-expert review loops
Non-technical subject-matter experts annotate traces and build datasets that feed back into evaluators and prompt optimisation.
Ideal For
Best For
- ✓Regression-testing multi-turn conversational and voice agents with simulated user personas before every release
- ✓Self-hosting LLM observability on Docker, Kubernetes or fully on-premises without buying an enterprise contract
- ✓Scoring live production traffic with online evaluators and clustering conversations by topic automatically
- ✓Automatic prompt optimisation using the DSPy approach instead of hand-tuning prompts by trial and error
- ✓EU-based teams needing ISO 27001, GDPR and EU data residency for their AI observability data
Not Ideal For
- ✗Teams that want a mature, heavily-reviewed incumbent — LangWatch is a EUR 1M pre-seed company with 3.5k GitHub stars and almost no G2, Capterra or Hacker News footprint to check its claims against
- ✗Organisations that need everything under a permissive licence; the core is Apache 2.0 but enterprise modules require a commercial licence, and the repository carries roughly 650 open issues
- ✗Buyers requiring a SOC 2 or HIPAA attestation, neither of which LangWatch publishes alongside its ISO 27001 certification
Integrations
Deployment
Market Analysis
Pros
- ✓Simulation testing with a user simulator and LLM judge closes a real gap, giving multi-turn agents CI-style regression tests rather than traces alone
- ✓Self-hosting on Docker, Kubernetes or on-premises is available on every plan including the free tier, which is rare in this category
- ✓OpenTelemetry-native and framework-agnostic across LangChain, LangGraph, CrewAI, Google ADK and n8n, which limits lock-in
- ✓Named enterprise references in regulated sectors — Deloitte, Backbase, PagBank and Visma — plus ISO 27001, GDPR and EU data residency
Cons
- ✗Very early company: a EUR 1 million pre-seed raised in February 2025 is a real continuity risk for a platform that sits in the production path
- ✗Open-core rather than fully open — enterprise modules require a commercial licence, so the Apache 2.0 headline does not cover everything
- ✗Roughly 650 open issues against only 3.5k GitHub stars suggests a maintenance backlog relative to the project's size
- ✗Almost no independent review footprint: no meaningful G2, Capterra, Product Hunt or Hacker News discussion exists to validate the vendor's claims
- ✗No SOC 2 or HIPAA attestation published, which will stall US healthcare and some enterprise procurement despite the ISO 27001 certificate
Pricing
Developer
$0
- ✓50K events per month
- ✓14-day data retention
- ✓2 users
- ✓3 scenarios, 3 simulations and 3 custom evals
- ✓Community support on GitHub and Discord
- ✓No credit card required
Growth
From $34/mo
- ✓EUR 29 per core seat per month
- ✓200K events per month included
- ✓Overage at EUR 5 per 100K events
- ✓30-day retention, then EUR 3 per GB
- ✓Unlimited lite users
- ✓Unlimited simulations, evals and prompts
- ✓Private Slack or Teams channel
Enterprise
Contact for pricing
- ✓Hybrid, self-hosted or on-premises deployment
- ✓SSO, RBAC and audit logs
- ✓Custom data retention and SLAs
- ✓Forward-deployed engineer and solution architect
- ✓AWS and Google Marketplace billing
- ✓ISO 27001 reports and custom DPA
Pricing is published in euros and metered on events rather than seats alone. The free Developer tier allows 50K events monthly with 14-day retention and two users; Growth is EUR 29 per core seat per month with 200K events included, EUR 5 per additional 100K events, and EUR 3 per GB for data retained beyond 30 days, with unlimited lite users and volume discounts above 20 users. Self-hosting is available on every tier including free, but SSO, RBAC, audit logs, custom retention and SLAs are Enterprise-only and unpriced.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 4 on domains other than langwatch.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Composio
Managed tool-calling and authentication layer connecting AI agents to 1,000+ enterprise applications
LangSmith
Framework-agnostic platform to observe, evaluate, deploy and continuously improve production AI agents
Kitesurf
An agent-first browser that runs in V8 isolates on Cloudflare Workers
Natural
Payments infrastructure that lets AI agents hold, send and collect money