Gimlet Cloud
by Gimlet Labs
Multi-silicon inference cloud that splits AI agent workloads across GPUs, CPUs and accelerators
Gimlet Cloud is a managed inference service from Gimlet Labs that breaks AI models and agent pipelines into stages and runs each stage on the chip best suited to it. It is aimed at frontier labs, hyperscalers and very large inference buyers who want more throughput and lower cost than a single-vendor GPU fleet gives them.
Gimlet Labs is a San Francisco company that came out of stealth in October 2025 and sells Gimlet Cloud, which it calls an agent-native, multi-silicon inference cloud. The approach is heterogeneous disaggregation: the platform splits a model's inference into phases such as prefill and decode, and an agent pipeline into model and non-model steps such as search, code sandboxes and MCP tool calls, then schedules each piece on the hardware that suits it. Gimlet works with silicon from NVIDIA, AMD, Intel, Arm, Cerebras and d-Matrix, and it does not make its own chip. Its published benchmarks include a B200 plus Intel Gaudi 3 prefill-decode split that it estimated at 1.7x better total cost of ownership, a d-Matrix Corsair speculative-decoding setup with 2-5x end-to-end speedups for interactive workloads, and a September 2026 Cerebras integration reaching up to 3,000 tokens per second. The company claims up to 10x throughput and interactivity gains for some workloads; those are its own numbers, not independent tests. It raised a $12M seed, an $80M Series A led by Menlo Ventures in March 2026, and a $300M Series B led by Andreessen Horowitz in September 2026 at a reported $3B valuation, for $392M in total. Gimlet says it has signed one of the top three frontier labs and one of the top three hyperscalers, has billions of dollars in contracted revenue, and is scaling to hundreds of megawatts of managed capacity. It also runs kforge, a multi-agent system that generates GPU kernels from PyTorch code. Access is sales-led: there is no published pricing, model list or self-serve signup.
Heads of AI infrastructure at labs, hyperscalers or large enterprises running inference at a scale where hardware utilisation decides the budget.
More tokens per dollar and per watt by matching each phase of a workload to a different chip type instead of buying one GPU family for everything.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Contact for pricing
- Target Market
- CTOs, Heads of AI Infrastructure, ML Platform Engineers, Hyperscalers
- Deployment
- Cloud-first, API-based
- Headquarters
- San Francisco, USA
- Customers
- Undisclosed; the company says it tripled its customer base in early 2026 and counts a top-three frontier lab and a top-three hyperscaler
Key Features
- ✓Multi-silicon disaggregation
Splits inference phases such as prefill and decode across different chip types so each phase runs on the hardware that serves it most efficiently.
- ✓Agent-native pipeline serving
Runs single agents and multi-agent systems end to end, including non-model stages like web search, code sandboxes, MCP servers and data sources.
- ✓Managed cloud or on-premises deployment
The same stack runs as Gimlet's managed cloud or as a managed service inside a customer's own data centre for large workloads.
- ✓Broad hardware support
Schedules work across NVIDIA, AMD, Intel, Arm, Cerebras and d-Matrix silicon, which lets buyers add accelerators without rewriting serving code.
- ✓Cerebras low-latency tier
A September 2026 integration adds Cerebras wafer-scale compute to Gimlet Cloud, reaching up to 3,000 tokens per second for interactive use.
- ✓kforge kernel generation
A multi-agent system that writes optimised CUDA, ROCm and Metal kernels from PyTorch code, with correctness checks and formal-verification research behind it.
Capabilities
Use Cases
- •Cutting inference TCO across mixed fleets
A platform team runs prefill on one vendor's GPUs and decode on another's accelerators, lowering cost per token versus a single-vendor fleet.
- •Low-latency agent responses
An agent product routes decode-heavy steps to Cerebras-class hardware so multi-step agent loops return answers faster to end users.
- •Serving multi-model agent pipelines
A team imports an existing agent pipeline that chains several models with search and code tools, and Gimlet schedules and optimises the whole graph.
- •Private large-scale inference
A hyperscaler or large enterprise deploys Gimlet's stack inside its own data centre to run heterogeneous hardware under a managed service.
Ideal For
Best For
- ✓Large inference buyers who want to reduce dependence on a single GPU vendor
- ✓Agent pipelines that mix model calls with search, code execution and tool calls
- ✓Latency-sensitive serving where SRAM-centric accelerators like Cerebras or d-Matrix help decode speed
- ✓Organisations that want a managed inference stack deployed inside their own data centre
Not Ideal For
- ✗Small and mid-sized teams that need self-serve signup and published per-token pricing; Gimlet is contact-sales only
- ✗Buyers who need documented compliance attestations up front, since no SOC 2 or similar certification is listed on its sites
- ✗Teams that want a public catalogue of hosted open models to compare against Together AI or Fireworks AI; Gimlet does not publish one
Deployment
Market & Ratings
Undisclosed; the company says it tripled its customer base in early 2026 and counts a top-three frontier lab and a top-three hyperscaler
Market Analysis
Pros
- ✓Strong financing and a reported $3B valuation lower vendor-viability risk for a young company
- ✓Publishes detailed engineering benchmarks on disaggregated inference and SRAM-centric chips
- ✓Works across six chip vendors, which gives buyers a path away from single-vendor GPU supply
- ✓Can run inside a customer's own data centre as well as in Gimlet's cloud
Cons
- ✗No public pricing, docs, model list or self-serve access; evaluation requires a sales engagement
- ✗Performance claims (up to 10x) come from the company's own benchmarks, and earlier kforge kernel results drew Hacker News criticism for loose correctness tolerances, weak baselines and no released code
- ✗No compliance certifications are listed on its websites, and named customers are not disclosed
- ✗Very young company: it emerged from stealth in October 2025, so production track record is short
Pricing
Gimlet Cloud
Contact for pricing
- ✓Managed inference API
- ✓Custom agent pipelines
- ✓Multi-silicon scheduling
Customer data centre deployment
Contact for pricing
- ✓Same stack deployed on customer infrastructure
- ✓Managed by Gimlet
- ✓For large-scale workloads
No list pricing is published. Gimlet's site routes both customers and hardware partners to a contact form, and its stated customers are frontier labs and hyperscalers signing large contracts, so expect negotiated capacity deals rather than per-token self-serve billing.
Security & Compliance
Sources
This page was written from 7 sources, 5 on domains other than gimletlabs.ai.
- 1.gimletlabs.ai — gimletlabs.aivendor
- 2.gimlet.ai — gimlet.ai
- 3.gimletlabs.ai — blogvendor
- 4.datacenter.news — gimlet labs raises usd 300m in andreessen led round
- 5.dealroom.co — 148837 gimlet labs raises 300m series b at 3b valuation to r
- 6.pulse2.com — gimlet labs raises 300 million series b led by andreessen ho
- 7.hn.algolia.com — 45118111
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
DigitalOcean Managed Agents
Managed agent runtime with microVM sandboxes, 16,000+ governed tools and serverless inference, billed on active CPU
Modular
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary
Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters