B

Bespoke Labs

by Bespoke Labs

Agent DevelopmentData & AnalyticsAI Models & APIsDeveloper Tools

The data and environments that train reliable AI agents

Free · Contact for pricing·Added Jul 8, 2026·Updated Sep 6, 2026
Share:
THE DAILY BRIEF
Bespoke Labs

by Bespoke Labs

Agent DevelopmentData & AnalyticsAI Models & APIsDeveloper Tools

The data and environments that train reliable AI agents

Free · Contact for pricing

Bespoke Labs is an applied AI research lab building the data curation and reinforcement-learning environment infrastructure used to train and evaluate long-horizon agents. It sells custom company-scale RL environments to frontier labs and enterprises, and publishes the underlying toolchain — Curator, Evalchemy, OpenThoughts and the MiniCheck and OpenThinker model families — as open source.

At a Glance

Category
Agent Development
Pricing
Free, Contact for pricing
Target Market
ML Engineers, Research Scientists, AI Platform Teams, CTOs, Data Scientists
Deployment
Open-source, Self-hosted, API-based
Founded
2024
Headquarters
Mountain View, California, United States
Team Size
11-50

Key Features

  • ✓Curator
  • ✓Multi-provider batch inference
  • ✓Company-scale RL environments
  • ✓GEPA policy optimiser
  • ✓Bespoke-MiniCheck grounded factuality models
  • ✓Open datasets and reproducible recipes
  • ✓Evalchemy

Capabilities

✓text generation
✗image generation
✗video generation
✗code generation
✓workflow automation
✓api access
✗audio generation
✓fine tuning
✓agent orchestration

Use Cases

  • •Building a post-training dataset
  • •Training an agent on realistic internal systems
  • •Cheap hallucination gating in RAG
  • •Distilling a smaller reasoning model
  • •Optimising agent prompts systematically

Ideal For

Best For

  • ✓Generating synthetic post-training datasets at scale with batch inference across multiple model providers
  • ✓Building company-scale RL environments that mirror real codebases, microservices, tickets and internal comms for agent training
  • ✓Distilling reasoning behaviour into smaller open checkpoints using published, reproducible recipes
  • ✓Grounded factuality checking in RAG pipelines using a small dedicated model rather than a frontier LLM call
  • ✓Automated prompt and policy optimisation via GEPA instead of hand-tuning agent prompts

Not Ideal For

  • ✗Teams wanting a finished, self-serve SaaS product — the commercial side is custom research-led delivery with no published pricing or signup
  • ✗Buyers who need a verifiable compliance posture today: an independent vendor profile records SOC 2 status as unknown and no publicly named customers
  • ✗Organisations that only consume models and never train or post-train them; almost all of the value here sits upstream of inference

Market Analysis

Open-sourceResearch-ledDeveloper-first

Pros

  • ✓Unusually verifiable for a research vendor — datasets, model checkpoints and training recipes are public on Hugging Face and GitHub, so claims can be reproduced rather than taken on trust
  • ✓Curator has real adoption as an open-source library (Apache-2.0, ~1.7k stars, 145 forks) with batch support across every major provider plus local backends
  • ✓OpenThoughts is reported at 500,000+ downloads with 190+ public models trained on it, and MiniCheck-7B topped the LLM-AggreFact factuality leaderboard
  • ✓Strong technical bench and investor signal: founders from Google DeepMind and UC Berkeley, with Jeff Dean and Anthropic/OpenAI/Meta operators on the cap table

Cons

  • ✗No publicly named customers — an independent vendor profile records the 'Fortune 500 enterprises and frontier labs' claim as vendor-asserted and unverified
  • ✗SOC 2 and other compliance certifications are not published; the same profile lists SOC 2 status as unknown, which is a real blocker for regulated buyers
  • ✗No pricing, no self-serve path and no product trial on the commercial side — evaluation requires a scoping conversation
  • ✗Curator carries 57 open issues against a small maintainer team, and practitioner discussion on Hacker News is thin (13 points on the launch thread) and largely posted by the founders, so independent production experience is hard to find
  • ✗Terminal-Bench, which the company lists among its work, is maintained in the harbor-framework GitHub organisation rather than under Bespoke Labs, so attribution is shared rather than exclusive
  • ✗At roughly 40-48 people it is a young company selling into a market where Scale AI and Surge AI have far deeper enterprise delivery capacity

Pricing

Open source (Curator, Evalchemy, datasets, models)

$0

  • ✓Apache-2.0 Curator library via pip
  • ✓Open datasets and model checkpoints on Hugging Face
  • ✓You pay only your own inference provider costs

Custom RL environments and data curation

Contact for pricing

  • ✓Company-scale RL environment build
  • ✓Custom evaluation and curation delivery
  • ✓Research-led engagement with frontier labs and enterprises

There is no published price list. The open-source side — Curator, Evalchemy, OpenThoughts, MiniCheck, Stratos and OpenThinker — is free under Apache-2.0 and costs only whatever you spend at your own inference provider, which for a large synthetic-data run is the dominant line item. The commercial side is custom delivery of RL environments and curation work sold directly to frontier labs and enterprises, quoted per engagement; an independent vendor profile classifies it as research-led custom delivery rather than a self-serve product, so expect a scoping conversation rather than a signup.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Bespoke Labs is an applied AI research lab building the data curation and reinforcement-learning environment infrastructure used to train and evaluate long-horizon agents. It sells custom company-scale RL environments to frontier labs and enterprises, and publishes the underlying toolchain — Curator, Evalchemy, OpenThoughts and the MiniCheck and OpenThinker model families — as open source.

Bespoke Labs is an applied AI research lab in Mountain View, California, founded in 2024 by Mahesh Sathiamoorthy (co-founder and CEO) and Alex Dimakis (co-founder and Chief Science Officer), that builds the data and environment infrastructure used to train and evaluate long-horizon AI agents. Its commercial work is company-scale reinforcement-learning environments: sandboxes replicating a real organisation's codebases, microservices, logs, tickets, email and Slack, so an agent can practise multi-step workflows against something closer to production than a toy benchmark. That sits on top of an unusually visible open-source stack. Curator (Apache-2.0, roughly 1.7k GitHub stars) is a Python library for synthetic data curation and structured extraction, with batch-mode support for OpenAI, Anthropic and Gemini plus LiteLLM, DeepSeek, kluster.ai and local vLLM or Ollama backends, fault recovery, Pydantic structured outputs, a hosted trace viewer, and fine-tuning handoff to Tinker and Fireworks AI. Evalchemy handles evaluation and benchmarking. OpenThoughts is an open reasoning dataset the company reports at over 500,000 downloads, used by Thinking Machines Lab, Meta and Amazon. Bespoke-MiniCheck-7B is a compact grounded-factuality checker — fine-tuned on roughly 35k examples — that topped the LLM-AggreFact leaderboard, and the Bespoke-Stratos and OpenThinker families distil reasoning behaviour into 7B and 32B checkpoints published on Hugging Face across 27 models and 23 datasets. GEPA, its genetic-Pareto optimizer, automates prompt and policy tuning and is claimed by the vendor to run in production at 200+ teams. In July 2026 it announced $40M across an 8VC-led seed and a Wing VC-led Series A, with Mayfield, The House Fund, Jeff Dean and Tristan Handy participating; headcount is roughly 40-48.

Ideal Buyer

ML platform and post-training teams at frontier labs or large enterprises that need realistic RL environments and curated training data to make long-horizon agents reliable.

Key Benefit

Agents trained and measured against sandboxes that mirror your real systems, instead of benchmarks that flatter them.

At a Glance

Category
Agent Development
Pricing
Free, Contact for pricing
Target Market
ML Engineers, Research Scientists, AI Platform Teams, CTOs, Data Scientists
Deployment
Open-source, Self-hosted, API-based
Founded
2024
Headquarters
Mountain View, California, United States
Team Size
11-50

Key Features

  • ✓
    Curator

    Apache-2.0 Python library for synthetic data pipelines with structured Pydantic outputs, async batching and fault recovery, so a long generation run survives provider failures.

  • ✓
    Multi-provider batch inference

    Native batch mode across OpenAI, Anthropic and Gemini plus LiteLLM, DeepSeek, vLLM and Ollama, letting a curation run use cheap local models and frontier APIs in one pipeline.

  • ✓
    Company-scale RL environments

    Sandboxes replicating real codebases, microservices, logs, tickets, email and Slack so agents practise long-horizon workflows against production-shaped systems.

  • ✓
    GEPA policy optimiser

    A genetic-Pareto optimiser that automates prompt and policy tuning, replacing manual prompt iteration with a measurable search over candidates.

  • ✓
    Bespoke-MiniCheck grounded factuality models

    Small fact-checking models that verify whether a claim is supported by its context, giving RAG systems a cheap hallucination gate instead of a frontier-model call.

  • ✓
    Open datasets and reproducible recipes

    OpenThoughts, OpenThinker and Bespoke-Stratos ship as public datasets and checkpoints on Hugging Face, so training claims can be independently reproduced.

  • ✓
    Evalchemy

    Evaluation and benchmarking tooling that measures agent capability consistently across model versions rather than through one-off ad hoc scripts.

Capabilities

✓text generation
✗image generation
✗video generation
✗code generation
✓workflow automation
✓api access
✗audio generation
✓fine tuning
✓agent orchestration

Use Cases

  • •
    Building a post-training dataset

    A model team uses Curator to generate and filter millions of structured reasoning traces across several providers with automatic recovery from API failures.

  • •
    Training an agent on realistic internal systems

    An enterprise commissions an RL environment mirroring its own services so agents learn multi-step workflows before touching production.

  • •
    Cheap hallucination gating in RAG

    A platform team runs Bespoke-MiniCheck over generated answers to flag claims unsupported by retrieved context at a fraction of a frontier-model call.

  • •
    Distilling a smaller reasoning model

    A team reproduces the OpenThoughts and OpenThinker recipes to distil reasoning behaviour into a 7B or 32B checkpoint they can host themselves.

  • •
    Optimising agent prompts systematically

    An engineering team applies GEPA to search prompt and policy variants against a scored objective instead of tuning prompts by intuition.

Ideal For

Best For

  • ✓Generating synthetic post-training datasets at scale with batch inference across multiple model providers
  • ✓Building company-scale RL environments that mirror real codebases, microservices, tickets and internal comms for agent training
  • ✓Distilling reasoning behaviour into smaller open checkpoints using published, reproducible recipes
  • ✓Grounded factuality checking in RAG pipelines using a small dedicated model rather than a frontier LLM call
  • ✓Automated prompt and policy optimisation via GEPA instead of hand-tuning agent prompts

Not Ideal For

  • ✗Teams wanting a finished, self-serve SaaS product — the commercial side is custom research-led delivery with no published pricing or signup
  • ✗Buyers who need a verifiable compliance posture today: an independent vendor profile records SOC 2 status as unknown and no publicly named customers
  • ✗Organisations that only consume models and never train or post-train them; almost all of the value here sits upstream of inference

Integrations

✓SDK Available
SDK:Python

Deployment

✓On-Premise

Market Analysis

Open-sourceResearch-ledDeveloper-first

Pros

  • ✓Unusually verifiable for a research vendor — datasets, model checkpoints and training recipes are public on Hugging Face and GitHub, so claims can be reproduced rather than taken on trust
  • ✓Curator has real adoption as an open-source library (Apache-2.0, ~1.7k stars, 145 forks) with batch support across every major provider plus local backends
  • ✓OpenThoughts is reported at 500,000+ downloads with 190+ public models trained on it, and MiniCheck-7B topped the LLM-AggreFact factuality leaderboard
  • ✓Strong technical bench and investor signal: founders from Google DeepMind and UC Berkeley, with Jeff Dean and Anthropic/OpenAI/Meta operators on the cap table

Cons

  • ✗No publicly named customers — an independent vendor profile records the 'Fortune 500 enterprises and frontier labs' claim as vendor-asserted and unverified
  • ✗SOC 2 and other compliance certifications are not published; the same profile lists SOC 2 status as unknown, which is a real blocker for regulated buyers
  • ✗No pricing, no self-serve path and no product trial on the commercial side — evaluation requires a scoping conversation
  • ✗Curator carries 57 open issues against a small maintainer team, and practitioner discussion on Hacker News is thin (13 points on the launch thread) and largely posted by the founders, so independent production experience is hard to find
  • ✗Terminal-Bench, which the company lists among its work, is maintained in the harbor-framework GitHub organisation rather than under Bespoke Labs, so attribution is shared rather than exclusive
  • ✗At roughly 40-48 people it is a young company selling into a market where Scale AI and Surge AI have far deeper enterprise delivery capacity

Pricing

Open source (Curator, Evalchemy, datasets, models)

$0

  • ✓Apache-2.0 Curator library via pip
  • ✓Open datasets and model checkpoints on Hugging Face
  • ✓You pay only your own inference provider costs

Custom RL environments and data curation

Contact for pricing

  • ✓Company-scale RL environment build
  • ✓Custom evaluation and curation delivery
  • ✓Research-led engagement with frontier labs and enterprises

There is no published price list. The open-source side — Curator, Evalchemy, OpenThoughts, MiniCheck, Stratos and OpenThinker — is free under Apache-2.0 and costs only whatever you spend at your own inference provider, which for a large synthetic-data run is the dominant line item. The commercial side is custom delivery of RL environments and curation work sold directly to frontier labs and enterprises, quoted per engagement; an independent vendor profile classifies it as research-led custom delivery rather than a self-serve product, so expect a scoping conversation rather than a signup.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

Connect

Sources

This page was written from 6 sources, 5 on domains other than bespokelabs.ai.

  1. 1.bespokelabs.ai — bespokelabs.aivendor
  2. 2.github.com — curator
  3. 3.huggingface.co — bespokelabs
  4. 4.pulse2.com — bespoke labs raises 40 million to build reliable ai agent tr
  5. 5.rl-list.com — bespoke labs
  6. 6.github.com — terminal bench
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe