Lambda
by Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters
Lambda is an AI cloud provider that rents NVIDIA GPU compute — on-demand B200 and H100 instances, 1-Click Clusters of 16 to 2,000+ GPUs, and single-tenant GB300 NVL72 superclusters. It is for AI labs, enterprises and ML teams that need dedicated training and inference capacity outside the hyperscalers.
Lambda, founded in 2012 by machine-learning engineers and led by co-founder and CEO Stephen Balaban, is a San Francisco-based 'neocloud' that builds and operates GPU infrastructure for AI training and inference. Its product line spans self-serve on-demand instances (NVIDIA B200, H100, A100, GH200 and older cards, from one to eight GPUs), 1-Click Clusters of 16 to 2,000+ B200 or H100 GPUs rented on terms from two weeks to a year, single-tenant Superclusters built on NVIDIA GB300 NVL72 with Quantum-2 InfiniBand, and private cloud deployments; its site also lists HGX B300 and Vera Rubin NVL72 systems. Lambda emphasises single-tenant, hardware-isolated 'caged' clusters, SOC 2 Type II certification, managed operations and co-engineering support, and it maintains Lambda Stack, its preconfigured deep-learning software bundle. The company raised a $480M Series D in February 2025 at roughly a $2.5B valuation, then more than $1.5B in a Series E led by TWG Global and the US Innovative Technology Fund in November 2025, and it says it serves tens of thousands of customers from researchers to hyperscalers. Reporting in 2026 describes Microsoft and Nvidia as anchor partners, more than $520M of revenue in the fiscal year to September 2025 alongside a $175M loss, a $1.5B GPU lease-back arrangement with Nvidia, and pre-IPO fundraising talks ahead of a planned listing in the second half of 2026. It competes with CoreWeave, Nebius, Crusoe and the hyperscalers' GPU offerings.
Heads of ML infrastructure at AI-native companies and enterprises training or serving large models who need reserved, single-tenant NVIDIA capacity with InfiniBand networking.
Access to current-generation NVIDIA GPUs — from a single on-demand instance to 2,000+ GPU clusters — with published per-GPU-hour pricing and no hyperscaler lock-in.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Heads of ML Infrastructure, AI Researchers, Enterprise Developers
- Deployment
- Cloud-only
- Founded
- 2012
- Headquarters
- San Francisco, United States
- Customers
- Tens of thousands (company claim, Nov 2025)
Key Features
- ✓On-demand GPU instances
Self-serve one-to-eight GPU instances on NVIDIA B200, H100, A100 and GH200, billed per GPU-hour without long-term commitment.
- ✓1-Click Clusters
Production-ready clusters of 16 to 2,000+ B200 or H100 GPUs rentable for two weeks to a year, for training runs without building infrastructure.
- ✓Superclusters
Single-tenant NVIDIA GB300 NVL72 systems with Quantum-2 InfiniBand for frontier-scale training, with managed operations and co-engineering.
- ✓Single-tenant isolation
Hardware-level isolation via caged clusters and SOC 2 Type II certification reduce multi-tenant risk for sensitive model weights and data.
- ✓Lambda Stack
A preconfigured deep-learning software bundle with drivers and frameworks, so teams start training without manual CUDA setup.
- ✓Current NVIDIA roadmap access
Offers HGX B300 and Vera Rubin NVL72 systems, giving customers early access to new GPU generations as NVIDIA ships them.
Capabilities
Use Cases
- •Frontier model pre-training
Reserve a single-tenant GB300 NVL72 supercluster with InfiniBand to train a large foundation model on dedicated hardware.
- •Fine-tuning on a fixed-term cluster
Spin up a 64-GPU H100 1-Click Cluster for a few weeks to post-train an open-weight model, then release the capacity.
- •Inference capacity outside hyperscalers
Serve production models on B200 instances to diversify GPU supply away from a single hyperscaler and control per-hour cost.
- •Research experimentation
Give ML researchers self-serve single-GPU instances with Lambda Stack preinstalled for fast prototyping and benchmark runs.
Ideal For
Best For
- ✓Foundation-model pre-training on dedicated, single-tenant GB300 or B200 clusters
- ✓Fine-tuning and post-training runs on 16 to 2,000+ GPU 1-Click Clusters for fixed terms
- ✓ML teams wanting self-serve on-demand H100/B200 instances with a preinstalled deep-learning stack
- ✓Enterprises seeking GPU capacity outside AWS, Azure and Google Cloud for cost or availability reasons
- ✓Security-sensitive training workloads that require hardware-isolated, caged infrastructure
Not Ideal For
- ✗Teams needing short, bursty multi-node jobs — practitioners report large interlinked H100 capacity has historically required long reservations rather than on-demand access
- ✗Organisations that want a full managed-services cloud (databases, serverless, broad PaaS) rather than GPU compute
- ✗Buyers who cannot accept counterparty risk from a loss-making, pre-IPO provider with concentrated anchor customers
- ✗Hobbyists or small projects on a tight budget, for whom marketplace GPU providers are often cheaper
Deployment
Market & Ratings
Tens of thousands (company claim, Nov 2025)
Market Analysis
Pros
- ✓Broad range of current NVIDIA hardware, from single GPUs to GB300 NVL72 superclusters
- ✓Transparent published pricing for on-demand and fixed-term cluster capacity
- ✓Well capitalised after a $1.5B+ Series E, with Microsoft and Nvidia as anchor partners
- ✓Single-tenant isolation and SOC 2 Type II for sensitive training workloads
Cons
- ✗Hacker News practitioners report large interlinked H100 capacity historically required long reservations, not on-demand access
- ✗Users report occasional breakage in the Lambda-managed Ubuntu/CUDA stack after upgrades
- ✗Loss-making ($175M loss on $520M+ revenue, FY to Sept 2025) with customer-concentration risk flagged by financial press
- ✗Narrow product scope compared with hyperscalers — GPU compute rather than a full cloud platform
Pricing
On-demand Instances
From $0.69/GPU/hr
- ✓B200 SXM6 $6.69-6.99/GPU/hr
- ✓H100 SXM $3.99-4.29/GPU/hr
- ✓A100 $1.99-2.79/GPU/hr
- ✓GH200 $2.29/GPU/hr
- ✓1-8 GPUs, self-serve
1-Click Clusters
From $5.54/GPU/hr
- ✓16 to 2,000+ GPUs
- ✓H100 $5.54-6.16/GPU/hr
- ✓B200 $8.87-9.86/GPU/hr
- ✓2-week to 1-year terms
Superclusters / Private Cloud
Contact for pricing
- ✓Single-tenant GB300 NVL72
- ✓Quantum-2 InfiniBand
- ✓Multi-year reservations
- ✓Managed operations
Lambda publishes per-GPU-hour list prices for on-demand instances (e.g. H100 SXM $3.99-4.29, B200 $6.69-6.99) and for 1-Click Clusters priced by term length from two weeks to one year (H100 $5.54-6.16, B200 $8.87-9.86). Reservations beyond one year, Superclusters and private cloud require a sales quote. Prices exclude sales tax/VAT; there is no free tier.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 3 on domains other than lambda.ai.
- 1.lambda.ai — lambda.aivendor
- 2.lambda.ai — pricingvendor
- 3.lambda.ai — lambda raises over 1.5b from twg global usit to build superivendor
- 4.w.media — lambda secures us 1 5 billion to expand u s ai factories
- 5.finance.yahoo.com — microsoft nvidia anchor lambdas ai 181306538
- 6.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Anyscale
Managed Ray platform for scaling AI data processing, training, inference and RL across thousands of GPUs on any cloud
Chroma
Open-source (Apache 2.0) vector and hybrid search database for AI, with a serverless Chroma Cloud on object storage
LiteLLM
Open-source AI gateway: 100+ LLM APIs behind one OpenAI-compatible endpoint, with cost tracking and guardrails
CIQ Fuzzball
Sovereign AI and HPC orchestration: train, fine-tune and serve models on infrastructure you control