SambaNova
by SambaNova Systems
Full-stack AI inference on custom dataflow chips — in the cloud, on-prem, or sovereign.
SambaNova is a full-stack AI inference platform built on its own Reconfigurable Dataflow Unit (RDU) chips, offering cloud, managed and on-premises deployment of frontier open models. It targets enterprises, sovereign AI operators and service providers that need high-throughput, secure inference without depending on GPU supply.
SambaNova Systems, founded in 2017 and headquartered in San Jose, California, builds a full-stack AI inference platform spanning custom silicon, systems, and software. Its Reconfigurable Dataflow Unit (RDU) chips use a dataflow architecture with a three-tier memory design that lets a single node hold and switch between multiple frontier-scale models — a structural advantage over GPU clusters for multi-model agentic serving. The current SN50 RDU is optimized for agentic inference (the company reports 435 output tokens/second on models such as gpt-oss-120b and DeepSeek), while the SN40-16 targets low-power inference. The platform ships in several forms: SambaCloud (hosted inference with OpenAI-compatible APIs), SambaManaged (managed inference), SambaRack (on-premises appliances), and SambaStack (the integrated chips-to-model stack), all coordinated by SambaOrchestrator for auto-scaling, load balancing and model management across data centers. It serves open frontier models including DeepSeek-V3.1 (671B), Meta's Llama series, OpenAI's gpt-oss-120b and MiniMax M2.7. On July 8, 2026 SambaNova completed the first close of a $1 billion Series F at an $11 billion post-money valuation, led by General Atlantic with participation from BlackRock, Intel Capital, T. Rowe Price, Capital Group, Vista Equity Partners, Battery Ventures, Qatar Investment Authority and others; it also named JPMorganChase as an inference infrastructure partner, deploying SN40 and SN50 systems for secure on-premises inference. Proceeds are earmarked for capacity expansion and deployments with enterprises, neo-clouds and sovereign AI customers across Australia, Europe and the UK.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, CIOs, Heads of Infrastructure, Enterprise Developers, Government & Public Sector, Cloud/Neo-cloud Operators
- Founded
- 2017
- Headquarters
- San Jose, USA
- Customers
- Enterprises, neo-clouds, sovereign AI operators (Australia, Europe, UK) and service providers; JPMorganChase is a named inference partner
Key Features
- ✓RDU dataflow chips (SN50, SN40-16)
Custom Reconfigurable Dataflow Units with a three-tier memory design; the SN50 is tuned for agentic inference and the SN40-16 for low-power serving.
- ✓Multi-model on one node
The memory architecture lets a single node hold and switch between multiple frontier-scale models instead of dedicating clusters per model.
- ✓OpenAI-compatible APIs
SambaCloud exposes OpenAI-compatible endpoints so existing application code can point at SambaNova with minimal change.
- ✓SambaOrchestrator
Manages workloads across data centers with auto-scaling, load balancing and model management.
- ✓On-prem and sovereign deployment
SambaRack and SambaStack deliver the same inference stack inside a customer's own data center or a sovereign cloud.
- ✓Frontier open-model support
Serves DeepSeek-V3.1 (671B), Meta Llama 4, OpenAI gpt-oss-120b, MiniMax M2.7 and other open models.
Capabilities
Use Cases
- •Secure on-prem enterprise inference
JPMorganChase deploys SN40 and SN50 systems on-premises to run AI inference on sensitive financial data inside its own perimeter.
- •Sovereign AI
Governments and regional operators in Australia, Europe and the UK run national AI infrastructure on SambaNova's stack.
- •Agentic inference at scale
Serve high-token-volume agent workloads that need to switch between several frontier models with predictable throughput.
Ideal For
Best For
- ✓High-throughput inference for agentic workloads without depending on GPU supply
- ✓On-premises inference where data cannot leave the enterprise perimeter
- ✓Sovereign AI deployments for governments and regional cloud operators
- ✓Serving multiple frontier open models from a single node
Integrations
Deployment
Market & Ratings
Enterprises, neo-clouds, sovereign AI operators (Australia, Europe, UK) and service providers; JPMorganChase is a named inference partner
Market Analysis
Pros
- ✓Owns the full stack from silicon up, giving real differentiation on throughput and power
- ✓One of very few inference vendors offering credible on-prem and sovereign deployment
- ✓Deep capitalization ($1B first close at $11B) and a marquee financial-services customer
Cons
- ✗Custom silicon means a smaller ecosystem and tooling surface than the CUDA/GPU world
- ✗Throughput figures such as 435 tokens/second are vendor-reported benchmarks
- ✗Little transparent pricing for enterprise or on-prem deployments
Pricing
SambaCloud
Contact for pricing
- ✓Hosted inference
- ✓OpenAI-compatible APIs
- ✓Frontier open models
SambaRack / SambaStack (on-prem)
Contact for pricing
- ✓On-premises RDU systems
- ✓Full chips-to-model stack
- ✓SambaOrchestrator
SambaCloud publishes plan details at cloud.sambanova.ai/plans/pricing; enterprise and on-premises deployments are quoted directly.
Sources
This page was written from 2 sources, 1 on domains other than sambanova.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Tsuga
Bring-your-own-cloud observability that keeps telemetry, and its cost, inside your own AWS account
OpenObserve
Open-source, Rust-based observability on object storage — logs, metrics, traces and LLM telemetry in one binary
Coralogix
AI-native observability that queries logs, metrics and traces in place — index-free, in your own S3 bucket
SkyPilot
Run and scale AI workloads across Kubernetes, Slurm and 25+ clouds from one interface