B

Braintrust

by Braintrust Data, Inc.

Developer ToolsAI Agents & OrchestrationData & Analytics

Evals, tracing and a prompt playground for teams shipping AI agents

Freemium · Subscription · Usage-based·Added Jul 3, 2026·Updated Sep 1, 2026
Share:
THE DAILY BRIEF
Braintrust

by Braintrust Data, Inc.

Developer ToolsAI Agents & OrchestrationData & Analytics

Evals, tracing and a prompt playground for teams shipping AI agents

Freemium · Subscription · Usage-based

Braintrust is an evaluation and observability platform for LLM applications and AI agents. Engineering and product teams use it to trace production agent behaviour span by span, score outputs against versioned datasets before a release ships, and compare prompts and models side by side in a playground — replacing eyeballed spot checks with a repeatable regression suite for non-deterministic software.

At a Glance

Category
Developer Tools
Pricing
Freemium, Subscription, Usage-based
Target Market
CTOs, Enterprise Developers, Data Scientists, AI/ML Engineers
Deployment
Cloud-first, Self-hosted, Hybrid, Multi-cloud
Headquarters
San Francisco, United States

Key Features

  • ✓Observe
  • ✓Evaluate
  • ✓Loop
  • ✓Brainstore
  • ✓Autoevals
  • ✓Prompt playground
  • ✓Self-hosted and hybrid data plane

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •Pre-release regression gate
  • •Agent debugging in production
  • •Model migration
  • •Cost and latency regression tracking
  • •Failure-to-dataset loop

Ideal For

Best For

  • ✓Gating prompt and model changes with scored evaluations before they reach production
  • ✓Debugging multi-step agent runs where the failure is buried three tool calls deep
  • ✓Turning real production failures into permanent labelled regression datasets
  • ✓Comparing model families on your own task instead of a public benchmark
  • ✓Running human annotation and review loops alongside automated scorers
  • ✓Keeping trace data inside your own cloud account for regulatory reasons

Not Ideal For

  • ✗Solo developers and cost-sensitive small teams — Pro is $249/month against LangSmith Plus at $99, and Starter's 14-day retention and 1 GB/month data cap run out quickly
  • ✗Teams that require an open-source, self-auditable stack: Braintrust is proprietary, where Langfuse is MIT-licensed and fully forkable
  • ✗Organisations that will not centralise LLM provider API keys with a third party — Braintrust disclosed unauthorised access to an AWS account holding exactly those keys in May 2026 and asked every customer to rotate
  • ✗Buyers who want a single flat price: metering across model credits, processed GB, scores and retention makes the bill hard to forecast

Market Analysis

Enterprise-gradeEvaluation-firstDeveloper-first

Pros

  • ✓Evaluation, tracing, datasets and playground in one product, so a failure found in production becomes a regression test without leaving the tool
  • ✓Real self-hosting and hybrid BYOC via Terraform on AWS, GCP and Azure — rare among eval vendors and the answer to EU data residency
  • ✓Named production users at Notion, Stripe, Cloudflare, Vercel, Ramp, Dropbox and Replit, so it is proven at genuine agent-trace volume
  • ✓SOC 2 Type II, GDPR, and HIPAA support via BAA on Enterprise
  • ✓Published list pricing with unlimited seats, so adding reviewers and PMs costs nothing

Cons

  • ✗Braintrust confirmed unauthorised access to an AWS account holding customer LLM provider API keys in May 2026 and told every customer to rotate; the company said there was no evidence of data theft, but the keys were centralised there to begin with
  • ✗Expensive against the field: $249/month Pro is a $150 premium over LangSmith Plus at $99, and a third-party comparison put 1M traces at roughly $101/month on Langfuse
  • ✗Proprietary and closed-source, so scoring internals cannot be audited or forked — the standard argument for choosing Langfuse instead
  • ✗Multi-dimensional metering across model credits, processed GB, scores and retention makes spend hard to forecast versus a flat per-unit price
  • ✗Thin independent practitioner discussion: the original Show HN drew 8 points and 2 comments, and G2 carries no readable rating, so hands-on write-ups are scarce relative to the funding

Pricing

Starter

$0

  • ✓$10/month model credits
  • ✓1 GB/month processed data
  • ✓10K scores/month
  • ✓14-day data retention
  • ✓Unlimited users, projects and experiments

Pro

From $249/mo

  • ✓$249/month model credits
  • ✓5 GB/month processed data
  • ✓50K scores/month
  • ✓30-day data retention
  • ✓Extended retention at $0.50/GB/month

Enterprise

Contact for pricing

  • ✓SAML SSO and custom RBAC roles
  • ✓DPA and HIPAA BAA
  • ✓S3 data export and custom retention
  • ✓Uptime SLA with shared Slack channel
  • ✓Custom dashboards and environments

List pricing is published, which is unusual in this category. Metering is multi-dimensional — model credits, GB of processed trace data, number of scores and retention window — rather than per seat, so users and projects are unlimited on every tier but a chatty agent can outrun a plan fast. Overage runs $4/GB on Starter and $3/GB on Pro, scores at $2.50 and $1.50 per 1,000 respectively, and extended retention $0.50/GB/month. SAML SSO, custom RBAC, DPA and HIPAA BAA, S3 export and an uptime SLA are all Enterprise-only, and neither Enterprise nor self-hosting carries a published price.

Security & Compliance

✓soc2
✓gdpr
✓hipaa
✗iso27001
✓sso
✓data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Braintrust is an evaluation and observability platform for LLM applications and AI agents. Engineering and product teams use it to trace production agent behaviour span by span, score outputs against versioned datasets before a release ships, and compare prompts and models side by side in a playground — replacing eyeballed spot checks with a repeatable regression suite for non-deterministic software.

Braintrust is an evaluation and observability platform for teams running LLM applications and agents in production, built by Braintrust Data, Inc. and used by engineering teams at Notion, Stripe, Cloudflare, Vercel, Ramp, Dropbox, Coursera, Replit and Box. The product is organised around four surfaces. Observe streams production traces — prompts, responses, tool calls and nested agent spans — in real time. Evaluate runs experiments against versioned datasets and scores outputs with code, LLM-as-judge or human graders, including the vendor's open-source Autoevals scorer library. Discover clusters production data into Topics that surface recurring failure modes without an engineer writing the query first. Loop is an in-product AI assistant that drafts improved prompts, scorers and datasets from a team's own observed traces. Underneath sits Brainstore, a database the company wrote specifically for deeply nested, high-cardinality trace data, on the argument that conventional observability stores choke on agent traces. Instrumentation is via SDKs in Python, TypeScript/JavaScript and Java plus an OpenTelemetry path, with the open-source braintrust-proxy fronting model providers. Deployment ranges from SaaS to full self-hosting, where the customer runs the entire data plane — API, PostgreSQL, Redis, object storage and Brainstore — inside their own AWS, GCP or Azure account using published Terraform modules while Braintrust retains only the control plane for UI, auth and metadata; a hybrid BYOC option has Braintrust operate that data plane inside the customer's cloud for EU residency. The company raised an $80M Series B led by Iconiq in February 2026 at an $800M valuation.

Ideal Buyer

The engineering lead who already has an LLM feature in production and cannot say whether last week's prompt change made it better or worse.

Key Benefit

A versioned, scored regression suite for non-deterministic output, so a prompt or model swap is gated by evidence rather than by vibes.

At a Glance

Category
Developer Tools
Pricing
Freemium, Subscription, Usage-based
Target Market
CTOs, Enterprise Developers, Data Scientists, AI/ML Engineers
Deployment
Cloud-first, Self-hosted, Hybrid, Multi-cloud
Headquarters
San Francisco, United States

Key Features

  • ✓
    Observe

    Streams production traces of prompts, responses and tool calls so a failed agent run can be replayed span by span.

  • ✓
    Evaluate

    Runs experiments against versioned datasets and scores them with code, LLM-as-judge or human graders.

  • ✓
    Loop

    An in-product AI assistant that drafts improved prompts, scorers and datasets from your own observed production traces.

  • ✓
    Brainstore

    A purpose-built trace database for deeply nested agent spans, with full-text search across millions of traces.

  • ✓
    Autoevals

    An open-source library of prebuilt scorers so teams do not hand-write factuality or relevance graders from scratch.

  • ✓
    Prompt playground

    Side-by-side prompt and model comparison across providers, shareable with non-engineers who own the copy.

  • ✓
    Self-hosted and hybrid data plane

    Terraform modules run the API, PostgreSQL, Redis and Brainstore inside your own AWS, GCP or Azure account.

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •
    Pre-release regression gate

    Block a deploy when a prompt change drops accuracy on the golden dataset below an agreed threshold.

  • •
    Agent debugging in production

    Replay a customer's failed support-agent conversation span by span to find which tool call returned bad context.

  • •
    Model migration

    Score a candidate model against the incumbent on your own eval set before switching providers or versions.

  • •
    Cost and latency regression tracking

    Watch token spend and response time per prompt version in real time to catch an expensive regression early.

  • •
    Failure-to-dataset loop

    Convert thumbs-down production traces into labelled test cases so the same failure cannot silently return.

Ideal For

Best For

  • ✓Gating prompt and model changes with scored evaluations before they reach production
  • ✓Debugging multi-step agent runs where the failure is buried three tool calls deep
  • ✓Turning real production failures into permanent labelled regression datasets
  • ✓Comparing model families on your own task instead of a public benchmark
  • ✓Running human annotation and review loops alongside automated scorers
  • ✓Keeping trace data inside your own cloud account for regulatory reasons

Not Ideal For

  • ✗Solo developers and cost-sensitive small teams — Pro is $249/month against LangSmith Plus at $99, and Starter's 14-day retention and 1 GB/month data cap run out quickly
  • ✗Teams that require an open-source, self-auditable stack: Braintrust is proprietary, where Langfuse is MIT-licensed and fully forkable
  • ✗Organisations that will not centralise LLM provider API keys with a third party — Braintrust disclosed unauthorised access to an AWS account holding exactly those keys in May 2026 and asked every customer to rotate
  • ✗Buyers who want a single flat price: metering across model credits, processed GB, scores and retention makes the bill hard to forecast

Integrations

✓SDK Available
SDK:PythonTypeScriptJavaScriptJava

Deployment

✓On-Premise

Market Analysis

Enterprise-gradeEvaluation-firstDeveloper-first

Pros

  • ✓Evaluation, tracing, datasets and playground in one product, so a failure found in production becomes a regression test without leaving the tool
  • ✓Real self-hosting and hybrid BYOC via Terraform on AWS, GCP and Azure — rare among eval vendors and the answer to EU data residency
  • ✓Named production users at Notion, Stripe, Cloudflare, Vercel, Ramp, Dropbox and Replit, so it is proven at genuine agent-trace volume
  • ✓SOC 2 Type II, GDPR, and HIPAA support via BAA on Enterprise
  • ✓Published list pricing with unlimited seats, so adding reviewers and PMs costs nothing

Cons

  • ✗Braintrust confirmed unauthorised access to an AWS account holding customer LLM provider API keys in May 2026 and told every customer to rotate; the company said there was no evidence of data theft, but the keys were centralised there to begin with
  • ✗Expensive against the field: $249/month Pro is a $150 premium over LangSmith Plus at $99, and a third-party comparison put 1M traces at roughly $101/month on Langfuse
  • ✗Proprietary and closed-source, so scoring internals cannot be audited or forked — the standard argument for choosing Langfuse instead
  • ✗Multi-dimensional metering across model credits, processed GB, scores and retention makes spend hard to forecast versus a flat per-unit price
  • ✗Thin independent practitioner discussion: the original Show HN drew 8 points and 2 comments, and G2 carries no readable rating, so hands-on write-ups are scarce relative to the funding

Pricing

Starter

$0

  • ✓$10/month model credits
  • ✓1 GB/month processed data
  • ✓10K scores/month
  • ✓14-day data retention
  • ✓Unlimited users, projects and experiments

Pro

From $249/mo

  • ✓$249/month model credits
  • ✓5 GB/month processed data
  • ✓50K scores/month
  • ✓30-day data retention
  • ✓Extended retention at $0.50/GB/month

Enterprise

Contact for pricing

  • ✓SAML SSO and custom RBAC roles
  • ✓DPA and HIPAA BAA
  • ✓S3 data export and custom retention
  • ✓Uptime SLA with shared Slack channel
  • ✓Custom dashboards and environments

List pricing is published, which is unusual in this category. Metering is multi-dimensional — model credits, GB of processed trace data, number of scores and retention window — rather than per seat, so users and projects are unlimited on every tier but a chatty agent can outrun a plan fast. Overage runs $4/GB on Starter and $3/GB on Pro, scores at $2.50 and $1.50 per 1,000 respectively, and extended retention $0.50/GB/month. SAML SSO, custom RBAC, DPA and HIPAA BAA, S3 export and an uptime SLA are all Enterprise-only, and neither Enterprise nor self-hosting carries a published price.

Security & Compliance

✓soc2
✓gdpr
✓hipaa
✗iso27001
✓sso
✓data residency

Connect

Sources

This page was written from 8 sources, 5 on domains other than braintrust.dev.

  1. 1.braintrust.dev — braintrust.devvendor
  2. 2.braintrust.dev — pricingvendor
  3. 3.braintrust.dev — self hostingvendor
  4. 4.siliconangle.com — braintrust lands 80m series b funding round become observabi
  5. 5.techcrunch.com — ai evaluation startup braintrust confirms breach tells every
  6. 6.dev.to — braintrust vs langsmith is 249mo worth it the may 2026 math
  7. 7.github.com — braintrustdata
  8. 8.hn.algolia.com — hn.algolia.com
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe