Respan
by Respan (formerly Keywords AI)
Observability, evals and an LLM gateway for AI agents in one control plane
Respan is an LLM engineering platform that combines agent tracing, automated and human evaluations, prompt management and a multi-provider AI gateway in a single product. It is aimed at engineering teams running agents in production who currently stitch together an observability tool, an eval framework and a routing proxy, and who want failures detected, root-caused and fixed in one loop.
Respan, which shipped as Keywords AI through 2025 and rebranded in early 2026, is a unified LLM engineering platform built around the argument that observability, evaluation and model routing should not be three separate vendors. The product has four surfaces: tracing that captures agent execution paths at 100% by default with no sampling, evaluations run either as LLM-as-judge or human review and executable online against live production traces rather than only offline, prompt versioning and optimisation, and the Respan Gateway, launched on 11 June 2026, which fronts more than 500 models behind one API endpoint with provider fallback. The differentiating piece is an automated evaluation agent that analyses failures across trials, identifies the root cause of specific decisions an agent made, and recommends which evaluation to add next, with alerts pushed to Slack, email or SMS; the company recommends using a different model family for evaluation than for execution to avoid self-grading bias, and low-temperature settings for determinism. Respan is model-, vendor-, framework- and language-agnostic and integrates in a few lines of code, running across development, staging and production. It was founded in 2023 by Hanhe Li (CEO) and Raymond Huang (CTO), went through Y Combinator's Winter 2024 batch, and operates from San Francisco with around 15 employees. In March 2026 it raised a $5M seed led by Gradient with Y Combinator, Hat-Trick Capital, Xiaoxiao Fund, Antigravity Capital and Alpen Capital participating. The company reports 100+ startup and enterprise customers, over 1 billion logs and 2 trillion tokens processed monthly, and more than 6.5 million end users served. Its core repository is Apache-2.0 licensed Python.
The engineering lead who already has an agent in production and is currently running Langfuse or LangSmith for traces, a separate eval harness, and LiteLLM or a hand-rolled proxy for routing
One control plane where a production failure is captured, automatically root-caused, and turned into the next evaluation, instead of three tools that never quite join up
At a Glance
- Category
- Developer Tools
- Pricing
- Freemium, Subscription, Contact for pricing
- Target Market
- AI Engineers, Enterprise Developers, ML Platform Leads, CTOs, Data Scientists
- Deployment
- Cloud-first, Self-hosted, API-based, Open-source
- Founded
- 2023
- Headquarters
- San Francisco, United States
- Team Size
- 11-50
- Customers
- 100+
Key Features
- ✓100% trace capture with no sampling
Every agent execution path is recorded by default, so the one failing run is never the one that was sampled away
- ✓Automated evaluation agent
Analyses failures across trials, root-causes specific agent decisions and recommends which evaluation to add next
- ✓Online evaluations on production traces
Evals run against live traffic rather than only a curated offline dataset, catching drift that offline suites miss
- ✓Respan Gateway
Single API endpoint fronting 500+ models with provider fallback, launched 11 June 2026, removing the need for a separate routing proxy
- ✓Prompt management and optimisation
Versioned prompts with optimisation tooling, so a prompt change is tracked and measurable rather than an untracked edit
- ✓LLM-as-judge and human review
Supports both automated grading and human-in-the-loop scoring, with guidance to use a different model family for judging to avoid bias
- ✓Alerting to Slack, email and SMS
Failures surface where the on-call engineer already is, rather than requiring someone to watch a dashboard
Capabilities
Use Cases
- •Debugging a non-deterministic agent failure
Replay the full captured execution path of a failing run to find which tool call or decision diverged from expectation
- •Catching quality regressions after a model swap
Run online evaluations against production traffic so a provider or version change surfaces as a measured score drop
- •Multi-provider failover
Route through one gateway endpoint across 500+ models so a provider outage degrades to a fallback instead of an incident
- •Building an evaluation suite from real failures
Let the evaluation agent root-cause production failures and recommend the next eval, growing coverage from real traffic
- •Controlling token spend across teams
Centralise traffic through the gateway to see per-team and per-feature token consumption against the 2T tokens processed monthly
Ideal For
Best For
- ✓Teams running LLM agents in production who need full-fidelity traces without sampling to debug non-deterministic failures
- ✓Consolidating a separate observability tool, evaluation framework and model-routing proxy into one platform with one integration
- ✓Running online evaluations against live production traffic rather than only offline against a curated dataset
- ✓Multi-provider model strategies that need fallback and a single endpoint across 500+ models without maintaining a proxy
- ✓Small engineering teams that want automated failure root-causing and eval recommendations instead of building an eval discipline from scratch
Not Ideal For
- ✗Organisations that require self-hosting: Respan restricts self-hosted deployment to the Enterprise tier, whereas Langfuse is MIT-licensed and fully self-hostable for free, which matters in compliance-sensitive environments
- ✗Teams whose primary need is rigorous offline evaluation with datasets and experiments — Respan's own comparison concedes it is less specialised there than Braintrust
- ✗Buyers who weight open-source community size heavily: the core repository sits at roughly 43 stars against Langfuse's far larger ecosystem and third-party integration set
- ✗Enterprise procurement processes that require an established vendor — this is a 15-person, seed-stage company holding your production LLM traffic in the gateway path
Integrations
Deployment
Market & Ratings
100+
Market Analysis
Pros
- ✓Published, self-serve pricing with a real free tier, which is rare in LLMOps and lets a team evaluate without a sales call
- ✓Consolidating traces, evals, prompts and gateway into one integration removes the plumbing most teams currently maintain between three vendors
- ✓Meaningful production scale for a seed-stage company: 1B+ logs and 2T+ tokens per month across 100+ customers and 6.5M+ end users
- ✓The automated evaluation agent turns production failures into eval coverage, which is the part teams reliably fail to do by hand
- ✓Core repository is Apache-2.0 licensed and actively maintained, with commits as recent as August 2026
Cons
- ✗Self-hosting is Enterprise-only, a clear disadvantage against MIT-licensed Langfuse for compliance-sensitive or air-gapped teams — a limitation Respan concedes in its own comparison
- ✗Less specialised than Braintrust for rigorous offline evaluation workflows built on datasets and experiments
- ✗The open-source footprint is small at roughly 43 stars and 11 forks, so community integrations and third-party support are thin next to Langfuse
- ✗A 15-person seed-stage vendor sitting in the gateway path of production LLM traffic is a concentration risk that enterprise procurement will question
- ✗The Keywords AI to Respan rebrand splits documentation, blog posts and third-party references across two names, making prior art harder to find
- ✗No G2, Capterra or TrustRadius listing and no Hacker News discussion, so there is no independent user feedback to check the vendor's claims against
Pricing
Free
$0
- ✓100k logs
- ✓1k scores
- ✓5 datasets
- ✓2 evaluators
- ✓5 prompts
- ✓No credit card required
Team
From $199/mo
- ✓Everything in Free
- ✓Unlimited datasets
- ✓Unlimited evaluators
- ✓Unlimited prompts
- ✓Private Slack channel
- ✓SOC 2 report
- ✓17% discount billed yearly
Enterprise
Contact for pricing
- ✓Everything in Team
- ✓Custom packages and volume discounts
- ✓Custom SLAs
- ✓Dedicated support engineer
- ✓HIPAA BAA
- ✓Self-hosting
Pricing is published, which is unusual in this category: a genuinely free tier capped at 100k logs, 1k scores, 5 datasets, 2 evaluators and 5 prompts, then a $199/month Team plan with 17% off annual billing that lifts those caps and includes the SOC 2 report. Self-hosting, HIPAA BAAs, custom SLAs and volume discounts are all gated behind a contact-for-pricing Enterprise tier.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 4 on domains other than respan.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Meta Muse Code
Meta's terminal coding agent for large repositories, with persistent background agents and the most aggressive token pricing in the category
Niteshift
The full-stack cloud for coding agents — real environments, verified pull requests
Opik
Open-source tracing, evaluation and guardrails for LLM applications and AI agents
Latitude
Open-source observability and evaluation for AI agents — find and fix failures before production