TrueFoundry
by TrueFoundry
Kubernetes-native AI gateway and deployment platform that runs inside your own VPC
TrueFoundry is an enterprise AI gateway and deployment platform that governs LLM, MCP tool and agent traffic from one control plane. It routes calls to 1000+ models across 30+ providers with fallbacks, budgets, guardrails and tracing, and unusually for this category it can run entirely inside the customer's own VPC, on-premises or air-gapped rather than only as vendor-hosted SaaS.
TrueFoundry sells the governance layer between enterprise applications and the model and tool providers they call. Its core is an AI Gateway - a proxy sitting in front of 1000+ models across 30+ providers including OpenAI, Anthropic, Google Vertex, AWS Bedrock, Azure, Cohere and Mistral - offering routing, automatic fallbacks, rate limits, budgets, caching, guardrails and OpenTelemetry tracing behind drop-in-compatible OpenAI and Anthropic SDK interfaces, so application code rarely changes to adopt it. Around that sit an MCP Gateway giving agents governed discovery of and access to tools, an Agent Gateway applying identity and access controls to agent-to-agent communication, and an AI deployment platform covering model serving, training and fine-tuning on Kubernetes. The company also maintains open-source work: Cognita, a RAG framework that reached 142 points on Hacker News in April 2024, AITori for intercepting AI application traffic, and TrueForge, a vendor-neutral agent harness. The distinguishing architectural choice is deployment flexibility: fully managed SaaS, a hybrid mode where TrueFoundry operates the gateway while LLM request and response data stays in the customer's own object storage, or fully self-hosted control and gateway planes in the customer's cloud, on-premises or air-gapped for strict data residency. Founded in 2021 by Nikunj Bajaj, Abhishek Choudhary and Anuraag Gutgutia, ex-Facebook engineers and UC Berkeley alumni, the company operates from San Francisco and Bengaluru with roughly 128 staff, and raised a $19M Series A in February 2025 led by Intel Capital with Eniac Ventures, Peak XV's Surge and Jump Capital, taking total funding to $21M. Its security page claims SOC 2, HIPAA and GDPR compliance, and it names Cargill, NetApp, NVIDIA, ResMed, Siemens Healthineers, Aviva and Automation Anywhere among customers.
The platform engineering team at a Kubernetes-run enterprise that must govern LLM, tool and agent traffic centrally and cannot let prompts and responses leave its own VPC - the constraint that rules out most hosted gateway vendors.
One governed proxy in front of 1000+ models and every MCP tool, with per-team budgets, guardrails, RBAC and audit logging, deployable inside your own network rather than a vendor's.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Subscription, Usage-based, Freemium, Contact for pricing
- Target Market
- CTOs, CIOs, Platform Engineers, MLOps Engineers, Enterprise Developers
- Deployment
- Hybrid, Self-hosted, Cloud-first, Multi-cloud
- Founded
- 2021
- Headquarters
- San Francisco, United States
- Team Size
- 51-200
Key Features
- ✓AI Gateway
A proxy across 1000+ models and 30+ providers with routing, automatic fallbacks, rate limits, caching and cost controls behind one endpoint.
- ✓MCP Gateway
Centralised discovery and access control for agent-to-tool connections, addressing the N-by-M integration sprawl of connecting every agent to every tool.
- ✓Agent Gateway
Applies identity, access control and policy enforcement to agent-to-agent communication so autonomous systems remain governed and attributable.
- ✓Flexible deployment planes
SaaS, hybrid with request data kept in your own object storage, or fully self-hosted control and gateway planes for strict data residency.
- ✓Observability and tracing
Framework-agnostic OpenTelemetry tracing with immutable audit logging and real-time policy enforcement across models, tools and agents.
- ✓Deployment and fine-tuning platform
Kubernetes-based model serving, training and fine-tuning, so the same platform covers custom models rather than only third-party APIs.
- ✓Native SDK compatibility
Drop-in support for the OpenAI and Anthropic SDKs, so existing application code points at the gateway without a rewrite.
Capabilities
Use Cases
- •Central LLM cost governance
A platform team sets per-team budgets and rate limits at the gateway so runaway experiments cannot silently consume a quarter's model spend.
- •Data-resident model access
A regulated enterprise self-hosts the gateway plane so prompts and responses never leave its own network while still reaching external providers.
- •Provider failover
Health-aware routing shifts traffic to a secondary provider automatically when a primary model endpoint degrades or exhausts its quota.
- •Governed MCP tool access
Agents discover and call internal tools through a single access-controlled gateway rather than each agent holding its own long-lived credentials.
- •Self-hosted model serving
A team fine-tunes and serves an open-weight model on its own GPU cluster while routing it through the same gateway as external providers.
Ideal For
Best For
- ✓Centralising LLM spend and rate limits across many teams so per-application provider keys stop being the unit of governance
- ✓Regulated or data-resident workloads that require the gateway plane to run inside the customer's own VPC, on-premises or air-gapped
- ✓Giving agents governed access to MCP tools through one discovery and access-control point instead of per-agent credentials
- ✓Provider redundancy - automatic fallback and health-aware routing when a model endpoint degrades or a quota is exhausted
- ✓Teams that also need model deployment and fine-tuning on Kubernetes rather than only a routing proxy
Not Ideal For
- ✗Small teams or prototypes: it is Kubernetes-native by design, so self-hosting assumes a cluster and someone to operate it, and open-source LiteLLM costs nothing to try
- ✗Buyers who want published, self-serve pricing at the low end - the first paid tier is $499/month, and self-hosting adds roughly $600-1,000/month of infrastructure on top
- ✗Single-provider shops committed to one hyperscaler, where Bedrock, Vertex AI or Microsoft Foundry already supply routing and governance inside an existing identity boundary
- ✗Procurement teams that need heavy third-party validation before signing - independent review coverage is thin and mostly gated
Integrations
Deployment
Market Analysis
Pros
- ✓One platform covers LLM gateway, MCP gateway, agent gateway, deployment and tracing - fewer vendors than assembling LiteLLM plus Portkey plus a Kubernetes serving stack
- ✓Runs fully inside the customer's own VPC, on-premises or air-gapped, which is the deciding factor for regulated buyers that hosted gateways cannot satisfy
- ✓SOC 2, HIPAA and GDPR compliance claimed on its security page, with RBAC, immutable audit logging and real-time policy enforcement
- ✓1000+ models across 30+ providers with drop-in OpenAI and Anthropic SDK compatibility, so migration rarely requires application changes
- ✓Backed by Intel Capital and Peak XV with named enterprise customers including Cargill, NetApp, NVIDIA, ResMed and Siemens Healthineers
Cons
- ✗Kubernetes-native by design - the self-hosted planes assume a cluster and an operator, which is real cost for a team that only wants a hosted proxy
- ✗The $499/month entry paid tier is a steep step from open-source LiteLLM at $0, and Pro caps at 1M requests and 10 users before $499-per-block overages
- ✗Self-hosting adds roughly $600-1,000/month of infrastructure on top of the licence, which is easy to miss when comparing sticker prices
- ✗Independent validation is thin: G2 and Gartner Peer Insights listings exist but block unattended access, Hacker News threads are mostly the company's own project launches at single-digit points, and most head-to-head gateway comparisons circulating online are published on TrueFoundry's own blog
- ✗The performance and cost figures (3-4 ms overhead, 350+ RPS per vCPU, 50% lower cloud spend, 80% higher GPU utilisation) are all vendor benchmarks with no third-party replication found
- ✗No ISO 27001 certification is listed on the security page, which some European procurement processes treat as mandatory
Pricing
Developer
$0
- ✓50,000 requests/month, 3 users, 10 saved prompts
- ✓Community support, no SLA, basic AI Gateway and prompt management
Pro
From $499/mo
- ✓1M requests/month, 10 users, unlimited saved prompts
- ✓Caching, control centre, observability, production support and standard SLA
Pro Plus
From $2,999/mo
- ✓1M requests/month, 25 users, advanced routing, guardrails, RBAC and SSO
- ✓Priority support with dedicated onboarding and an enterprise-grade SLA
Enterprise
Contact for pricing
- ✓10M+ requests/month, unlimited users, full customisation
- ✓VPC, on-premises and air-gapped deployment for control and gateway planes
Priced on requests processed and active platform users together, not on tokens. The free Developer tier covers 50,000 requests a month for three users; Pro is $499/month for 1M requests and 10 users, with overage sold as an extra 2M requests and 5 API keys for a further $499; Pro Plus is $2,999/month for 25 users and is where advanced routing, guardrails, RBAC and SSO unlock; Enterprise is quote-only at 10M+ requests with unlimited users. All tiers carry a 7-day free trial. Hosting cost is separate and often overlooked: managed SaaS adds none, but self-hosting the gateway or both planes runs roughly $600-1,000/month of infrastructure on top of the licence.
Security & Compliance
Connect
Sources
This page was written from 7 sources, 3 on domains other than truefoundry.com.
- 1.truefoundry.com — truefoundry.comvendor
- 2.truefoundry.com — pricingvendor
- 3.truefoundry.com — securityvendor
- 4.truefoundry.com — docsvendor
- 5.hn.algolia.com — search
- 6.cloudnuro.ai — ai gateway buyers guide 2026 kong vs portkey vs litellm vs t
- 7.intelcapital.com — truefoundry secures 19 million series a funding to transform
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Gimlet Cloud
Multi-silicon inference cloud that splits AI agent workloads across GPUs, CPUs and accelerators
DigitalOcean Managed Agents
Managed agent runtime with microVM sandboxes, 16,000+ governed tools and serverless inference, billed on active CPU
Modular
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary