Bespoke Labs
by Bespoke Labs
The data and environments that train reliable AI agents
Bespoke Labs is an applied AI research lab building the data curation and reinforcement-learning environment infrastructure used to train and evaluate long-horizon agents. It sells custom company-scale RL environments to frontier labs and enterprises, and publishes the underlying toolchain — Curator, Evalchemy, OpenThoughts and the MiniCheck and OpenThinker model families — as open source.
Bespoke Labs is an applied AI research lab in Mountain View, California, founded in 2024 by Mahesh Sathiamoorthy (co-founder and CEO) and Alex Dimakis (co-founder and Chief Science Officer), that builds the data and environment infrastructure used to train and evaluate long-horizon AI agents. Its commercial work is company-scale reinforcement-learning environments: sandboxes replicating a real organisation's codebases, microservices, logs, tickets, email and Slack, so an agent can practise multi-step workflows against something closer to production than a toy benchmark. That sits on top of an unusually visible open-source stack. Curator (Apache-2.0, roughly 1.7k GitHub stars) is a Python library for synthetic data curation and structured extraction, with batch-mode support for OpenAI, Anthropic and Gemini plus LiteLLM, DeepSeek, kluster.ai and local vLLM or Ollama backends, fault recovery, Pydantic structured outputs, a hosted trace viewer, and fine-tuning handoff to Tinker and Fireworks AI. Evalchemy handles evaluation and benchmarking. OpenThoughts is an open reasoning dataset the company reports at over 500,000 downloads, used by Thinking Machines Lab, Meta and Amazon. Bespoke-MiniCheck-7B is a compact grounded-factuality checker — fine-tuned on roughly 35k examples — that topped the LLM-AggreFact leaderboard, and the Bespoke-Stratos and OpenThinker families distil reasoning behaviour into 7B and 32B checkpoints published on Hugging Face across 27 models and 23 datasets. GEPA, its genetic-Pareto optimizer, automates prompt and policy tuning and is claimed by the vendor to run in production at 200+ teams. In July 2026 it announced $40M across an 8VC-led seed and a Wing VC-led Series A, with Mayfield, The House Fund, Jeff Dean and Tristan Handy participating; headcount is roughly 40-48.
ML platform and post-training teams at frontier labs or large enterprises that need realistic RL environments and curated training data to make long-horizon agents reliable.
Agents trained and measured against sandboxes that mirror your real systems, instead of benchmarks that flatter them.
At a Glance
- Category
- Agent Development
- Pricing
- Free, Contact for pricing
- Target Market
- ML Engineers, Research Scientists, AI Platform Teams, CTOs, Data Scientists
- Deployment
- Open-source, Self-hosted, API-based
- Founded
- 2024
- Headquarters
- Mountain View, California, United States
- Team Size
- 11-50
Key Features
- ✓Curator
Apache-2.0 Python library for synthetic data pipelines with structured Pydantic outputs, async batching and fault recovery, so a long generation run survives provider failures.
- ✓Multi-provider batch inference
Native batch mode across OpenAI, Anthropic and Gemini plus LiteLLM, DeepSeek, vLLM and Ollama, letting a curation run use cheap local models and frontier APIs in one pipeline.
- ✓Company-scale RL environments
Sandboxes replicating real codebases, microservices, logs, tickets, email and Slack so agents practise long-horizon workflows against production-shaped systems.
- ✓GEPA policy optimiser
A genetic-Pareto optimiser that automates prompt and policy tuning, replacing manual prompt iteration with a measurable search over candidates.
- ✓Bespoke-MiniCheck grounded factuality models
Small fact-checking models that verify whether a claim is supported by its context, giving RAG systems a cheap hallucination gate instead of a frontier-model call.
- ✓Open datasets and reproducible recipes
OpenThoughts, OpenThinker and Bespoke-Stratos ship as public datasets and checkpoints on Hugging Face, so training claims can be independently reproduced.
- ✓Evalchemy
Evaluation and benchmarking tooling that measures agent capability consistently across model versions rather than through one-off ad hoc scripts.
Capabilities
Use Cases
- •Building a post-training dataset
A model team uses Curator to generate and filter millions of structured reasoning traces across several providers with automatic recovery from API failures.
- •Training an agent on realistic internal systems
An enterprise commissions an RL environment mirroring its own services so agents learn multi-step workflows before touching production.
- •Cheap hallucination gating in RAG
A platform team runs Bespoke-MiniCheck over generated answers to flag claims unsupported by retrieved context at a fraction of a frontier-model call.
- •Distilling a smaller reasoning model
A team reproduces the OpenThoughts and OpenThinker recipes to distil reasoning behaviour into a 7B or 32B checkpoint they can host themselves.
- •Optimising agent prompts systematically
An engineering team applies GEPA to search prompt and policy variants against a scored objective instead of tuning prompts by intuition.
Ideal For
Best For
- ✓Generating synthetic post-training datasets at scale with batch inference across multiple model providers
- ✓Building company-scale RL environments that mirror real codebases, microservices, tickets and internal comms for agent training
- ✓Distilling reasoning behaviour into smaller open checkpoints using published, reproducible recipes
- ✓Grounded factuality checking in RAG pipelines using a small dedicated model rather than a frontier LLM call
- ✓Automated prompt and policy optimisation via GEPA instead of hand-tuning agent prompts
Not Ideal For
- ✗Teams wanting a finished, self-serve SaaS product — the commercial side is custom research-led delivery with no published pricing or signup
- ✗Buyers who need a verifiable compliance posture today: an independent vendor profile records SOC 2 status as unknown and no publicly named customers
- ✗Organisations that only consume models and never train or post-train them; almost all of the value here sits upstream of inference
Integrations
Deployment
Market Analysis
Pros
- ✓Unusually verifiable for a research vendor — datasets, model checkpoints and training recipes are public on Hugging Face and GitHub, so claims can be reproduced rather than taken on trust
- ✓Curator has real adoption as an open-source library (Apache-2.0, ~1.7k stars, 145 forks) with batch support across every major provider plus local backends
- ✓OpenThoughts is reported at 500,000+ downloads with 190+ public models trained on it, and MiniCheck-7B topped the LLM-AggreFact factuality leaderboard
- ✓Strong technical bench and investor signal: founders from Google DeepMind and UC Berkeley, with Jeff Dean and Anthropic/OpenAI/Meta operators on the cap table
Cons
- ✗No publicly named customers — an independent vendor profile records the 'Fortune 500 enterprises and frontier labs' claim as vendor-asserted and unverified
- ✗SOC 2 and other compliance certifications are not published; the same profile lists SOC 2 status as unknown, which is a real blocker for regulated buyers
- ✗No pricing, no self-serve path and no product trial on the commercial side — evaluation requires a scoping conversation
- ✗Curator carries 57 open issues against a small maintainer team, and practitioner discussion on Hacker News is thin (13 points on the launch thread) and largely posted by the founders, so independent production experience is hard to find
- ✗Terminal-Bench, which the company lists among its work, is maintained in the harbor-framework GitHub organisation rather than under Bespoke Labs, so attribution is shared rather than exclusive
- ✗At roughly 40-48 people it is a young company selling into a market where Scale AI and Surge AI have far deeper enterprise delivery capacity
Pricing
Open source (Curator, Evalchemy, datasets, models)
$0
- ✓Apache-2.0 Curator library via pip
- ✓Open datasets and model checkpoints on Hugging Face
- ✓You pay only your own inference provider costs
Custom RL environments and data curation
Contact for pricing
- ✓Company-scale RL environment build
- ✓Custom evaluation and curation delivery
- ✓Research-led engagement with frontier labs and enterprises
There is no published price list. The open-source side — Curator, Evalchemy, OpenThoughts, MiniCheck, Stratos and OpenThinker — is free under Apache-2.0 and costs only whatever you spend at your own inference provider, which for a large synthetic-data run is the dominant line item. The commercial side is custom delivery of RL environments and curation work sold directly to frontier labs and enterprises, quoted per engagement; an independent vendor profile classifies it as research-led custom delivery rather than a self-serve product, so expect a scoping conversation rather than a signup.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 5 on domains other than bespokelabs.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Mastra
Open-source TypeScript framework and platform for building, deploying and observing production AI agents and workflows
Agentrys
Agentic design automation that lets chip teams build and own a self-improving engineering workforce
OpenAI Agents API
The Codex agent harness behind one API call — managed sessions, sandboxes and subagents
Letta
Open-source platform for stateful AI agents with persistent memory, now with a TypeScript Agents SDK for embedding them in your own apps