Braintrust
by Braintrust Data, Inc.
Evals, tracing and a prompt playground for teams shipping AI agents
Braintrust is an evaluation and observability platform for LLM applications and AI agents. Engineering and product teams use it to trace production agent behaviour span by span, score outputs against versioned datasets before a release ships, and compare prompts and models side by side in a playground — replacing eyeballed spot checks with a repeatable regression suite for non-deterministic software.
Braintrust is an evaluation and observability platform for teams running LLM applications and agents in production, built by Braintrust Data, Inc. and used by engineering teams at Notion, Stripe, Cloudflare, Vercel, Ramp, Dropbox, Coursera, Replit and Box. The product is organised around four surfaces. Observe streams production traces — prompts, responses, tool calls and nested agent spans — in real time. Evaluate runs experiments against versioned datasets and scores outputs with code, LLM-as-judge or human graders, including the vendor's open-source Autoevals scorer library. Discover clusters production data into Topics that surface recurring failure modes without an engineer writing the query first. Loop is an in-product AI assistant that drafts improved prompts, scorers and datasets from a team's own observed traces. Underneath sits Brainstore, a database the company wrote specifically for deeply nested, high-cardinality trace data, on the argument that conventional observability stores choke on agent traces. Instrumentation is via SDKs in Python, TypeScript/JavaScript and Java plus an OpenTelemetry path, with the open-source braintrust-proxy fronting model providers. Deployment ranges from SaaS to full self-hosting, where the customer runs the entire data plane — API, PostgreSQL, Redis, object storage and Brainstore — inside their own AWS, GCP or Azure account using published Terraform modules while Braintrust retains only the control plane for UI, auth and metadata; a hybrid BYOC option has Braintrust operate that data plane inside the customer's cloud for EU residency. The company raised an $80M Series B led by Iconiq in February 2026 at an $800M valuation.
The engineering lead who already has an LLM feature in production and cannot say whether last week's prompt change made it better or worse.
A versioned, scored regression suite for non-deterministic output, so a prompt or model swap is gated by evidence rather than by vibes.
At a Glance
- Category
- Developer Tools
- Pricing
- Freemium, Subscription, Usage-based
- Target Market
- CTOs, Enterprise Developers, Data Scientists, AI/ML Engineers
- Deployment
- Cloud-first, Self-hosted, Hybrid, Multi-cloud
- Headquarters
- San Francisco, United States
Key Features
- ✓Observe
Streams production traces of prompts, responses and tool calls so a failed agent run can be replayed span by span.
- ✓Evaluate
Runs experiments against versioned datasets and scores them with code, LLM-as-judge or human graders.
- ✓Loop
An in-product AI assistant that drafts improved prompts, scorers and datasets from your own observed production traces.
- ✓Brainstore
A purpose-built trace database for deeply nested agent spans, with full-text search across millions of traces.
- ✓Autoevals
An open-source library of prebuilt scorers so teams do not hand-write factuality or relevance graders from scratch.
- ✓Prompt playground
Side-by-side prompt and model comparison across providers, shareable with non-engineers who own the copy.
- ✓Self-hosted and hybrid data plane
Terraform modules run the API, PostgreSQL, Redis and Brainstore inside your own AWS, GCP or Azure account.
Capabilities
Use Cases
- •Pre-release regression gate
Block a deploy when a prompt change drops accuracy on the golden dataset below an agreed threshold.
- •Agent debugging in production
Replay a customer's failed support-agent conversation span by span to find which tool call returned bad context.
- •Model migration
Score a candidate model against the incumbent on your own eval set before switching providers or versions.
- •Cost and latency regression tracking
Watch token spend and response time per prompt version in real time to catch an expensive regression early.
- •Failure-to-dataset loop
Convert thumbs-down production traces into labelled test cases so the same failure cannot silently return.
Ideal For
Best For
- ✓Gating prompt and model changes with scored evaluations before they reach production
- ✓Debugging multi-step agent runs where the failure is buried three tool calls deep
- ✓Turning real production failures into permanent labelled regression datasets
- ✓Comparing model families on your own task instead of a public benchmark
- ✓Running human annotation and review loops alongside automated scorers
- ✓Keeping trace data inside your own cloud account for regulatory reasons
Not Ideal For
- ✗Solo developers and cost-sensitive small teams — Pro is $249/month against LangSmith Plus at $99, and Starter's 14-day retention and 1 GB/month data cap run out quickly
- ✗Teams that require an open-source, self-auditable stack: Braintrust is proprietary, where Langfuse is MIT-licensed and fully forkable
- ✗Organisations that will not centralise LLM provider API keys with a third party — Braintrust disclosed unauthorised access to an AWS account holding exactly those keys in May 2026 and asked every customer to rotate
- ✗Buyers who want a single flat price: metering across model credits, processed GB, scores and retention makes the bill hard to forecast
Integrations
Deployment
Market Analysis
Pros
- ✓Evaluation, tracing, datasets and playground in one product, so a failure found in production becomes a regression test without leaving the tool
- ✓Real self-hosting and hybrid BYOC via Terraform on AWS, GCP and Azure — rare among eval vendors and the answer to EU data residency
- ✓Named production users at Notion, Stripe, Cloudflare, Vercel, Ramp, Dropbox and Replit, so it is proven at genuine agent-trace volume
- ✓SOC 2 Type II, GDPR, and HIPAA support via BAA on Enterprise
- ✓Published list pricing with unlimited seats, so adding reviewers and PMs costs nothing
Cons
- ✗Braintrust confirmed unauthorised access to an AWS account holding customer LLM provider API keys in May 2026 and told every customer to rotate; the company said there was no evidence of data theft, but the keys were centralised there to begin with
- ✗Expensive against the field: $249/month Pro is a $150 premium over LangSmith Plus at $99, and a third-party comparison put 1M traces at roughly $101/month on Langfuse
- ✗Proprietary and closed-source, so scoring internals cannot be audited or forked — the standard argument for choosing Langfuse instead
- ✗Multi-dimensional metering across model credits, processed GB, scores and retention makes spend hard to forecast versus a flat per-unit price
- ✗Thin independent practitioner discussion: the original Show HN drew 8 points and 2 comments, and G2 carries no readable rating, so hands-on write-ups are scarce relative to the funding
Pricing
Starter
$0
- ✓$10/month model credits
- ✓1 GB/month processed data
- ✓10K scores/month
- ✓14-day data retention
- ✓Unlimited users, projects and experiments
Pro
From $249/mo
- ✓$249/month model credits
- ✓5 GB/month processed data
- ✓50K scores/month
- ✓30-day data retention
- ✓Extended retention at $0.50/GB/month
Enterprise
Contact for pricing
- ✓SAML SSO and custom RBAC roles
- ✓DPA and HIPAA BAA
- ✓S3 data export and custom retention
- ✓Uptime SLA with shared Slack channel
- ✓Custom dashboards and environments
List pricing is published, which is unusual in this category. Metering is multi-dimensional — model credits, GB of processed trace data, number of scores and retention window — rather than per seat, so users and projects are unlimited on every tier but a chatty agent can outrun a plan fast. Overage runs $4/GB on Starter and $3/GB on Pro, scores at $2.50 and $1.50 per 1,000 respectively, and extended retention $0.50/GB/month. SAML SSO, custom RBAC, DPA and HIPAA BAA, S3 export and an uptime SLA are all Enterprise-only, and neither Enterprise nor self-hosting carries a published price.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 5 on domains other than braintrust.dev.
- 1.braintrust.dev — braintrust.devvendor
- 2.braintrust.dev — pricingvendor
- 3.braintrust.dev — self hostingvendor
- 4.siliconangle.com — braintrust lands 80m series b funding round become observabi
- 5.techcrunch.com — ai evaluation startup braintrust confirms breach tells every
- 6.dev.to — braintrust vs langsmith is 249mo worth it the may 2026 math
- 7.github.com — braintrustdata
- 8.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Momentic
Agentic QA platform that writes, runs and self-heals end-to-end tests for web and mobile apps
Raindrop
AI agent monitoring that catches silent production failures: Sentry for AI agents
GitLab Duo Agent Platform
Agentic AI across the whole GitLab DevSecOps lifecycle: planning, coding, code review, CI/CD and security agents under one governance model
CodeRabbit
AI code review and agentic change management for teams shipping human- and machine-written code