L

LangWatch

by LangWatch

Agent DevelopmentDeveloper ToolsGovernance & Security

Open-source agent testing platform pairing simulation-based QA with OpenTelemetry tracing and evals

Freemium · Subscription · Usage-based · Contact for pricing·Added Aug 16, 2026·Updated Aug 16, 2026
Share:
THE DAILY BRIEF
LangWatch

by LangWatch

Agent DevelopmentDeveloper ToolsGovernance & Security

Open-source agent testing platform pairing simulation-based QA with OpenTelemetry tracing and evals

Freemium · Subscription · Usage-based · Contact for pricing

LangWatch is an Apache-2.0 licensed LLMOps platform for teams putting AI agents into production. It combines OpenTelemetry-native tracing, online and offline evaluation, and end-to-end agent simulation — running scripted user personas against your agent before release — so quality problems are caught in CI rather than by customers.

At a Glance

Category
Agent Development
Pricing
Freemium, Subscription, Usage-based, Contact for pricing
Target Market
CTOs, AI Engineers, QA and Product Leads, Enterprise Developers
Deployment
Open-source, Self-hosted, Cloud-first, Hybrid
Headquarters
Amsterdam, Netherlands

Key Features

  • Agent simulation testing
  • OpenTelemetry-native tracing
  • Online and offline evaluation
  • DSPy-based prompt optimisation
  • AI gateway
  • Self-hosting on every tier
  • Langy AI engineer

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Pre-release agent QA
  • Turning incidents into tests
  • Regulated deployment in the EU
  • Cost and quality monitoring
  • Domain-expert review loops

Ideal For

Best For

  • Regression-testing multi-turn conversational and voice agents with simulated user personas before every release
  • Self-hosting LLM observability on Docker, Kubernetes or fully on-premises without buying an enterprise contract
  • Scoring live production traffic with online evaluators and clustering conversations by topic automatically
  • Automatic prompt optimisation using the DSPy approach instead of hand-tuning prompts by trial and error
  • EU-based teams needing ISO 27001, GDPR and EU data residency for their AI observability data

Not Ideal For

  • Teams that want a mature, heavily-reviewed incumbent — LangWatch is a EUR 1M pre-seed company with 3.5k GitHub stars and almost no G2, Capterra or Hacker News footprint to check its claims against
  • Organisations that need everything under a permissive licence; the core is Apache 2.0 but enterprise modules require a commercial licence, and the repository carries roughly 650 open issues
  • Buyers requiring a SOC 2 or HIPAA attestation, neither of which LangWatch publishes alongside its ISO 27001 certification

Market Analysis

Open-sourceDeveloper-firstEU-based

Pros

  • Simulation testing with a user simulator and LLM judge closes a real gap, giving multi-turn agents CI-style regression tests rather than traces alone
  • Self-hosting on Docker, Kubernetes or on-premises is available on every plan including the free tier, which is rare in this category
  • OpenTelemetry-native and framework-agnostic across LangChain, LangGraph, CrewAI, Google ADK and n8n, which limits lock-in
  • Named enterprise references in regulated sectors — Deloitte, Backbase, PagBank and Visma — plus ISO 27001, GDPR and EU data residency

Cons

  • Very early company: a EUR 1 million pre-seed raised in February 2025 is a real continuity risk for a platform that sits in the production path
  • Open-core rather than fully open — enterprise modules require a commercial licence, so the Apache 2.0 headline does not cover everything
  • Roughly 650 open issues against only 3.5k GitHub stars suggests a maintenance backlog relative to the project's size
  • Almost no independent review footprint: no meaningful G2, Capterra, Product Hunt or Hacker News discussion exists to validate the vendor's claims
  • No SOC 2 or HIPAA attestation published, which will stall US healthcare and some enterprise procurement despite the ISO 27001 certificate

Pricing

Developer

$0

  • 50K events per month
  • 14-day data retention
  • 2 users
  • 3 scenarios, 3 simulations and 3 custom evals
  • Community support on GitHub and Discord
  • No credit card required

Growth

From $34/mo

  • EUR 29 per core seat per month
  • 200K events per month included
  • Overage at EUR 5 per 100K events
  • 30-day retention, then EUR 3 per GB
  • Unlimited lite users
  • Unlimited simulations, evals and prompts
  • Private Slack or Teams channel

Enterprise

Contact for pricing

  • Hybrid, self-hosted or on-premises deployment
  • SSO, RBAC and audit logs
  • Custom data retention and SLAs
  • Forward-deployed engineer and solution architect
  • AWS and Google Marketplace billing
  • ISO 27001 reports and custom DPA

Pricing is published in euros and metered on events rather than seats alone. The free Developer tier allows 50K events monthly with 14-day retention and two users; Growth is EUR 29 per core seat per month with 200K events included, EUR 5 per additional 100K events, and EUR 3 per GB for data retained beyond 30 days, with unlimited lite users and volume discounts above 20 users. Self-hosting is available on every tier including free, but SSO, RBAC, audit logs, custom retention and SLAs are Enterprise-only and unpriced.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

LangWatch is an Apache-2.0 licensed LLMOps platform for teams putting AI agents into production. It combines OpenTelemetry-native tracing, online and offline evaluation, and end-to-end agent simulation — running scripted user personas against your agent before release — so quality problems are caught in CI rather than by customers.

LangWatch is an Amsterdam-based LLMOps platform founded by Manouk Draisma and Rogerio Chaves that positions itself less as a dashboard and more as a testing loop for AI agents. Its distinguishing capability is simulation: a scenario defines the agent under test, a user simulator that plays a realistic persona across text or voice, and an LLM judge that scores the resulting conversation — so a multi-turn agent can be regression-tested in CI the way deterministic software is, and any observed production failure can be converted into a simulation that verifies the fix and gates the release. Around that sit OpenTelemetry-native tracing over OTLP following the GenAI semantic conventions, dataset creation and manual annotation tools for domain experts, online evaluation scoring live production traffic, automatic prompt optimisation drawing on Stanford's DSPy framework, and an AI gateway for governance and cost control that the project measures at roughly 700 nanoseconds of hot-path overhead. The platform is framework-agnostic, with documented support for LangChain, LangGraph, CrewAI, Google ADK, Mastra, the Vercel AI SDK, LangFlow, Flowise and n8n, and model providers including OpenAI, Anthropic, Azure OpenAI, Vertex AI, Bedrock, Groq and Ollama. It is open-core: the core is Apache 2.0 with 3.5k GitHub stars, the SDKs are MIT-licensed, and enterprise modules require a commercial licence. Self-hosting via Docker Compose, Kubernetes Helm or full on-premises deployment is available on every tier, which is unusual in this category. LangWatch is ISO 27001 certified and GDPR compliant, names Deloitte, Backbase, PagBank, Visma and Freeday as customers, and raised a EUR 1 million pre-seed led by Passion Capital in February 2025.

Ideal Buyer

The AI engineering or QA lead at a regulated European enterprise who needs multi-turn agents regression-tested before release and must keep trace data inside their own infrastructure.

Key Benefit

Agent regressions get caught by simulated users in CI before release, instead of being discovered in production traces after customers have already hit them.

At a Glance

Category
Agent Development
Pricing
Freemium, Subscription, Usage-based, Contact for pricing
Target Market
CTOs, AI Engineers, QA and Product Leads, Enterprise Developers
Deployment
Open-source, Self-hosted, Cloud-first, Hybrid
Headquarters
Amsterdam, Netherlands

Key Features

  • Agent simulation testing

    Scripted user personas run realistic text or voice scenarios against your agent before it ever reaches production.

  • OpenTelemetry-native tracing

    Follows GenAI semantic conventions over OTLP, so instrumentation stays portable rather than locked to one vendor.

  • Online and offline evaluation

    LLM judges, custom code and workflow evaluators score both curated datasets and live production traffic.

  • DSPy-based prompt optimisation

    Automatically searches for better prompts using Stanford's DSPy approach instead of manual trial and error.

  • AI gateway

    OpenAI- and Anthropic-compatible proxy adding governance and cost control at roughly 700 nanoseconds hot-path overhead.

  • Self-hosting on every tier

    Docker Compose, Kubernetes Helm, on-premises and hybrid deployment are available without an enterprise contract.

  • Langy AI engineer

    Turns product goals into test plans and scenarios, scores them with judge rubrics and opens pull requests with fixes.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Pre-release agent QA

    Run a suite of simulated customer conversations against a new agent build and block the release if scores drop.

  • Turning incidents into tests

    Convert a real production failure trace into a repeatable simulation that proves the fix and stops the regression returning.

  • Regulated deployment in the EU

    Self-host the data plane on-premises under ISO 27001 and GDPR while using the managed control plane for the interface.

  • Cost and quality monitoring

    Track token spend per conversation alongside evaluator scores to see where quality drops as cheaper models are substituted.

  • Domain-expert review loops

    Non-technical subject-matter experts annotate traces and build datasets that feed back into evaluators and prompt optimisation.

Ideal For

Best For

  • Regression-testing multi-turn conversational and voice agents with simulated user personas before every release
  • Self-hosting LLM observability on Docker, Kubernetes or fully on-premises without buying an enterprise contract
  • Scoring live production traffic with online evaluators and clustering conversations by topic automatically
  • Automatic prompt optimisation using the DSPy approach instead of hand-tuning prompts by trial and error
  • EU-based teams needing ISO 27001, GDPR and EU data residency for their AI observability data

Not Ideal For

  • Teams that want a mature, heavily-reviewed incumbent — LangWatch is a EUR 1M pre-seed company with 3.5k GitHub stars and almost no G2, Capterra or Hacker News footprint to check its claims against
  • Organisations that need everything under a permissive licence; the core is Apache 2.0 but enterprise modules require a commercial licence, and the repository carries roughly 650 open issues
  • Buyers requiring a SOC 2 or HIPAA attestation, neither of which LangWatch publishes alongside its ISO 27001 certification

Integrations

SDK Available
SDK:PythonTypeScript

Deployment

On-Premise

Market Analysis

Open-sourceDeveloper-firstEU-based

Pros

  • Simulation testing with a user simulator and LLM judge closes a real gap, giving multi-turn agents CI-style regression tests rather than traces alone
  • Self-hosting on Docker, Kubernetes or on-premises is available on every plan including the free tier, which is rare in this category
  • OpenTelemetry-native and framework-agnostic across LangChain, LangGraph, CrewAI, Google ADK and n8n, which limits lock-in
  • Named enterprise references in regulated sectors — Deloitte, Backbase, PagBank and Visma — plus ISO 27001, GDPR and EU data residency

Cons

  • Very early company: a EUR 1 million pre-seed raised in February 2025 is a real continuity risk for a platform that sits in the production path
  • Open-core rather than fully open — enterprise modules require a commercial licence, so the Apache 2.0 headline does not cover everything
  • Roughly 650 open issues against only 3.5k GitHub stars suggests a maintenance backlog relative to the project's size
  • Almost no independent review footprint: no meaningful G2, Capterra, Product Hunt or Hacker News discussion exists to validate the vendor's claims
  • No SOC 2 or HIPAA attestation published, which will stall US healthcare and some enterprise procurement despite the ISO 27001 certificate

Pricing

Developer

$0

  • 50K events per month
  • 14-day data retention
  • 2 users
  • 3 scenarios, 3 simulations and 3 custom evals
  • Community support on GitHub and Discord
  • No credit card required

Growth

From $34/mo

  • EUR 29 per core seat per month
  • 200K events per month included
  • Overage at EUR 5 per 100K events
  • 30-day retention, then EUR 3 per GB
  • Unlimited lite users
  • Unlimited simulations, evals and prompts
  • Private Slack or Teams channel

Enterprise

Contact for pricing

  • Hybrid, self-hosted or on-premises deployment
  • SSO, RBAC and audit logs
  • Custom data retention and SLAs
  • Forward-deployed engineer and solution architect
  • AWS and Google Marketplace billing
  • ISO 27001 reports and custom DPA

Pricing is published in euros and metered on events rather than seats alone. The free Developer tier allows 50K events monthly with 14-day retention and two users; Growth is EUR 29 per core seat per month with 200K events included, EUR 5 per additional 100K events, and EUR 3 per GB for data retained beyond 30 days, with unlimited lite users and volume discounts above 20 users. Self-hosting is available on every tier including free, but SSO, RBAC, audit logs, custom retention and SLAs are Enterprise-only and unpriced.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Connect

Sources

This page was written from 6 sources, 4 on domains other than langwatch.ai.

  1. 1.langwatch.ailangwatch.aivendor
  2. 2.langwatch.aipricingvendor
  3. 3.github.comlangwatch
  4. 4.marktechpost.comlangwatch open sources the missing evaluation layer for ai a
  5. 5.siliconcanals.comlangwatch raises e1m in pre seed round
  6. 6.posthog.combest ai observability tools
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe