Cerebras
by Cerebras Systems Inc.
Wafer-scale AI silicon and the fastest inference cloud for open-weight models
Cerebras Systems builds the Wafer-Scale Engine, an AI processor occupying an entire silicon wafer, and sells it both as on-premises CS-3 systems and as a public inference cloud with drop-in OpenAI API compatibility. It targets teams whose applications are bottlenecked by token generation speed rather than model quality — real-time agents, coding assistants and long reasoning chains.
Cerebras Systems designs the Wafer-Scale Engine (WSE), an AI processor built on an entire silicon wafer rather than a die cut from one, which the company markets as "58x larger and 15x faster than GPUs." Founded in 2015 in Sunnyvale by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie and Jean-Philippe Fricker, it monetises that chip three ways: CS-3 systems installed in customer data centres, managed supercomputers, and the Cerebras Inference cloud, which offers drop-in OpenAI API compatibility and serves open-weight models including Llama, Qwen, GLM, Kimi and gpt-oss. The architectural bet is memory locality — because model weights sit in on-wafer SRAM rather than off-chip HBM, the memory-bandwidth wall that caps GPU token generation largely disappears. Cerebras has published 969 tokens per second on Llama 3.1 405B and roughly 1,500 tokens per second on Qwen3-235B with full 131k context support, each of which drew 350+ points on Hacker News. Cerebras Code packages the same speed as a flat coding subscription at $50 and $200 per month. The company is now public: it listed on Nasdaq as CBRS on 14 May 2026 at $185 per share, raising $5.55 billion at a $56.4 billion fully diluted valuation — the largest US technology IPO since Snowflake in 2020, following a $1.1B Series G in September 2025 and a $1B Series H in February 2026. Its S-1 shows revenue of $510.0M in 2025, up 76% year over year, on a 39% gross margin and a non-GAAP loss of $75.7M once a one-time $363.3M non-cash gain is stripped out. The same filing discloses that two UAE-affiliated customers — MBZUAI at 62% and G42 at 24% — produced 86% of 2025 revenue.
Platform or applied-AI engineering leaders whose product is latency-bound — voice agents, coding assistants, multi-step agent loops — and who have already accepted open-weight models.
Roughly an order of magnitude more tokens per second than GPU inference on the same open model, reachable by changing one base URL.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Subscription, Usage-based, Contact for pricing
- Target Market
- CTOs, Platform Engineering Leaders, Data Scientists, Enterprise Developers, HPC and Research Computing Leads
- Deployment
- API-based, Cloud-first, Self-hosted, Hybrid
- Founded
- 2015
- Headquarters
- Sunnyvale, California, United States
Key Features
- ✓Wafer-Scale Engine (WSE)
A processor built on a full silicon wafer, marketed as 58x larger than a GPU, which removes the multi-chip interconnect tax entirely.
- ✓Cerebras Inference cloud
Drop-in OpenAI API compatibility means existing SDK code can be repointed by changing the base URL rather than rewriting the integration.
- ✓On-wafer SRAM weight residency
Model weights stay on chip instead of in HBM, which is the specific reason published throughput reaches four figures per second.
- ✓CS-3 systems
On-premises appliances for organisations that must keep training and inference inside their own data centre boundary.
- ✓Cerebras Code subscriptions
Flat-rate coding plans at $50 and $200 per month carrying 24 million and 120 million daily token allowances with no weekly caps.
- ✓Open-model catalogue
Serves the Llama, Qwen, GLM, Kimi and gpt-oss families rather than one proprietary model, so weights stay portable.
Capabilities
Use Cases
- •Real-time coding assistant
Run Qwen3-Coder 480B at up to 2,000 tokens per second inside Cursor, Cline, Continue.dev or RooCode with no proprietary IDE lock-in.
- •Multi-step agent orchestration
Cut wall-clock time on agent loops where a dozen sequential model calls each add seconds of user-visible latency.
- •On-premises model training
Install CS-3 systems in a national lab or regulated data centre to train large models without renting external GPU capacity.
- •High-volume document reasoning
Process long-context summarisation and extraction jobs where output-token throughput, not model quality, is the binding constraint.
- •Voice and conversational AI
Serve sub-second responses in live voice agents, where typical GPU token rates make natural turn-taking difficult to achieve.
Ideal For
Best For
- ✓Latency-critical agent loops where every turn waits on token generation
- ✓Interactive coding assistants that need sub-second completions on large models
- ✓Reasoning and chain-of-thought workloads whose wall-clock cost is dominated by output tokens
- ✓Teams standardised on open-weight models who want a drop-in OpenAI-compatible endpoint
- ✓National labs and research institutions training frontier models on dedicated CS-3 clusters
Not Ideal For
- ✗Teams committed to proprietary frontier models — Cerebras serves open weights only, so there is no path to GPT-5 or Claude on this hardware
- ✗Buyers who want a full managed AI platform with vector stores, evaluation tooling and fine-tuning UIs; this is an inference endpoint and a chip, not a platform
- ✗Procurement functions with strict supplier-concentration or geopolitical-exposure policies, given the S-1 discloses 86% of 2025 revenue from two UAE-affiliated entities and TSMC as sole manufacturer with no long-term supply agreement
- ✗Regulated buyers who need published SOC 2 or ISO 27001 attestations before signing — Cerebras publishes no reachable public trust centre
Integrations
Deployment
Market Analysis
Pros
- ✓Token generation speed is genuinely differentiated rather than incremental — 969 tokens/s on Llama 3.1 405B is roughly an order of magnitude above GPU inference
- ✓Drop-in OpenAI API compatibility makes evaluation nearly free: repoint a base URL instead of rewriting an integration
- ✓Cerebras Code offers flat monthly pricing with large daily token allowances and explicitly no weekly caps, unusual among coding subscriptions
- ✓Revenue grew 76% to $510.0M in 2025 and the company is now publicly listed, which materially reduces counterparty risk on multi-year deals
Cons
- ✗Extreme customer concentration: the S-1 discloses MBZUAI at 62% and G42 at 24% — 86% of 2025 revenue from two UAE-affiliated entities, with MBZUAI at 77.9% of year-end accounts receivable
- ✗Gross margin is 39% and falling from 42%, low for a company valued as a semiconductor platform, and cloud services margin is only 30%
- ✗The S-1 discloses material weaknesses in accounting controls over revenue recognition, and no long-term supply agreement with TSMC, its sole manufacturer
- ✗Model choice is limited to open weights, so buyers wanting frontier proprietary models cannot consolidate onto this platform
- ✗Independent evidence is thin: the top Hacker News threads are all Cerebras' own blog posts and press releases, with little third-party benchmarking or production war-stories, and no reachable public trust centre for security attestations
Pricing
Free trial
$0
- ✓$5 in free credits
- ✓Access to all Cerebras-powered models
- ✓Community Discord support
Developer
From $10
- ✓Self-serve payment
- ✓10x higher rate limits than the free tier
- ✓Higher priority processing
Cerebras Code Pro
$50/mo
- ✓Up to 24 million tokens per day
- ✓Qwen3-Coder 480B
- ✓131k-token context window
- ✓Works with any OpenAI-compatible IDE
Cerebras Code Max
$200/mo
- ✓Up to 120 million tokens per day
- ✓No weekly limits
- ✓No proprietary IDE lock-in
Enterprise
Contact for pricing
- ✓Dedicated queue priority for lowest latency
- ✓Support for custom model weights
- ✓Model fine-tuning and training services
- ✓Dedicated support team with response-time guarantees
Cerebras publishes subscription prices but not per-token rates: a $5 free-credit trial, a self-serve Developer tier from $10 carrying 10x the free-tier rate limits, and Enterprise by quote for dedicated queue priority, custom model weights and fine-tuning services. The coding product is the only clearly published list price — Cerebras Code Pro at $50/month for up to 24 million tokens per day and Code Max at $200/month for up to 120 million, both with no weekly caps. CS-3 hardware and managed supercomputers are sold entirely through sales with no list price, so any capacity purchase requires a procurement cycle.
Security & Compliance
Sources
This page was written from 6 sources, 2 on domains other than cerebras.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
WEKA NeuralMesh
Microservices storage and memory fabric built for AI training and inference at exabyte scale
Tsuga
Bring-your-own-cloud observability that keeps telemetry, and its cost, inside your own AWS account
OpenObserve
Open-source, Rust-based observability on object storage — logs, metrics, traces and LLM telemetry in one binary
Coralogix
AI-native observability that queries logs, metrics and traces in place — index-free, in your own S3 bucket