Modular
by Modular (a Qualcomm company)
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
Modular builds MAX, an inference serving framework, and Mojo, a Python-like systems language for GPU kernels, plus a managed Modular Cloud. It is for platform and ML infrastructure teams that want to serve open models on several hardware vendors from one stack instead of being tied to CUDA.
Modular was founded in 2022 by Chris Lattner, the creator of LLVM and Swift, and Tim Davis, and is based in Palo Alto. Its stack has three parts. MAX is a model serving and modeling framework with OpenAI-compatible endpoints that ships as a container under 700 MB and runs on NVIDIA and AMD GPUs and x86 and ARM CPUs; the company also lists TPU, Trainium, Intel, Qualcomm and Apple silicon support. Mojo is a Python-styled systems language for writing portable GPU kernels; Mojo 1.0 shipped on August 11, 2026 in the Modular 26.5 release, and the compiler was open-sourced under Apache 2.0 with LLVM exceptions on August 18. Modular Cloud runs the same stack as shared per-token endpoints, dedicated per-GPU-minute endpoints, or inside a customer's own VPC. Modular raised a $250 million Series C at a $1.6 billion valuation in September 2025, led by Thomas Tull's USIT fund, bringing total funding to $380 million. Qualcomm completed its acquisition on July 29, 2026 in an all-stock deal reported at $3.9 billion, and says Mojo, MAX and Modular Cloud continue as products and that MAX will keep supporting third-party hardware. Named users include Inworld, MiniMax, Hippocratic AI, TensorWave and AWS. The vendor claims up to 2x the throughput of vLLM on the same hardware; those figures have not been independently verified. The main GitHub repository has about 29,900 stars and 3,200 forks.
A head of ML platform running open-weight models at scale who wants to buy AMD or other non-NVIDIA accelerators without rewriting the serving stack.
One inference stack and one kernel language that runs across GPU vendors, with self-hosting free and a managed cloud for burst capacity.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Free, Usage-based, Contact for pricing
- Target Market
- CTOs, ML Platform Teams, AI Infrastructure Engineers, Enterprise Developers
- Deployment
- Open-source, Self-hosted, Multi-cloud, Hybrid
- Founded
- 2022
- Headquarters
- Palo Alto, United States
- Team Size
- 51-200
Key Features
- ✓MAX inference server
Serves 1,000+ pre-configured models behind OpenAI-compatible REST endpoints from a container under 700 MB.
- ✓Mojo 1.0 language
A Python-styled systems language for writing GPU kernels once and running them across chip vendors, now stable at 1.0.
- ✓Open-source compiler
The Mojo compiler, standard library and MAX kernels are on GitHub under Apache 2.0 with LLVM exceptions, so teams can audit and fork them.
- ✓Multi-vendor hardware support
Runs on NVIDIA and AMD GPUs and x86 and ARM CPUs, with TPU, Trainium, Intel, Qualcomm and Apple silicon also listed, which reduces lock-in to one accelerator.
- ✓Modular Cloud endpoints
Shared per-token endpoints for testing and dedicated per-GPU-minute endpoints for production, with SOC 2 Type 2 certification.
- ✓Bring Your Own Cloud
Deploys the stack in the customer's VPC on their cloud credits, so data never leaves their environment.
- ✓Forward-deployed engineers
Cloud and BYOC plans include Modular engineers who tune throughput and latency for the customer's specific pipelines.
Capabilities
Use Cases
- •Cutting inference cost on AMD
Move open-model serving onto AMD GPUs without rewriting kernels; TensorWave reports 70% total cost savings on AMD with MAX.
- •Low-latency voice AI
Serve speech and LLM models with sub-second time to first token, as Inworld and Hippocratic AI do for real-time conversations.
- •Writing portable custom kernels
Write performance-critical GPU kernels in Mojo once and run them on several vendors instead of maintaining CUDA and ROCm versions.
- •Private inference in your VPC
Run the managed stack inside your own AWS, GCP, Azure or Oracle account to satisfy data residency and compliance policies.
Ideal For
Best For
- ✓Serving open-weight LLMs on AMD GPUs alongside or instead of NVIDIA
- ✓Teams writing custom GPU kernels who want one portable language instead of CUDA and ROCm
- ✓Running inference inside your own VPC with vendor engineers tuning the deployment
- ✓Low-latency voice and chat workloads that need sub-second time to first token
- ✓Free self-hosted inference with OpenAI-compatible endpoints
Not Ideal For
- ✗Teams that need hardware neutrality guaranteed for years, since Qualcomm now owns the stack and analysts flag a risk that its own accelerators get optimisation priority
- ✗Python shops expecting Mojo to run existing Python code; Hacker News practitioners note source compatibility with Python is not close
- ✗Buyers who need published enterprise pricing; BYOC and Enterprise are quote-only
Integrations
Deployment
Market Analysis
Pros
- ✓Genuine multi-vendor GPU support reduces dependence on NVIDIA supply and pricing
- ✓Open-source compiler and kernels with an active community of roughly 200 contributors on the 1.0 release
- ✓Free self-hosted edition with OpenAI-compatible endpoints lowers the cost of a trial
- ✓Named production users including Inworld, MiniMax and Hippocratic AI
Cons
- ✗Qualcomm ownership raises questions about long-term hardware neutrality and future pricing
- ✗Mojo looks like Python but does not run Python source, and some Hacker News users said they never got past toy examples
- ✗Performance claims against vLLM and CUDA are vendor-reported and lack independent verification
- ✗Licensing scope for MAX outside the open-sourced components was still unclear after the August 2026 release, per RuntimeWire
Pricing
Self-Hosted
$0
- ✓MAX and Mojo under community license
- ✓Deploy anywhere
- ✓Community support via Discord and GitHub
Modular Cloud
Usage-based (e.g. Qwen 3.5 9B $0.17/$0.25 per 1M tokens)
- ✓Shared endpoints billed per token
- ✓Dedicated endpoints billed per GPU minute
- ✓SOC 2 Type 2
- ✓Forward-deployed engineers
Bring Your Own Cloud
Contact for pricing
- ✓Billed per minute of reserved GPU capacity
- ✓Runs in your VPC
- ✓Custom APIs
Enterprise
Contact for pricing
- ✓AWS, GCP, Azure or Oracle
- ✓Hybrid deployment
- ✓Custom engagement
Self-hosting MAX and Mojo is free. Modular Cloud publishes per-token rates for shared endpoints (DeepSeek V4 at $1.74 input / $3.48 output per 1M tokens, Qwen 3.5 9B at $0.17 / $0.25) and bills dedicated endpoints per GPU minute. BYOC and Enterprise are quote-only, and one analysis warns enterprise pricing could change after the Qualcomm integration. Prices checked on modular.com/pricing on October 2, 2026.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 5 on domains other than modular.com.
- 1.modular.com — modular.comvendor
- 2.modular.com — pricingvendor
- 3.modular.com — modular 26 5 mojo 1 0 is herevendor
- 4.github.com — modular
- 5.sdxcentral.com — modular raises 250m for ais unified compute layer at 16b val
- 6.runtimewire.com — chris lattner open sources mojo qualcomm modular
- 7.developersdigest.tech — qualcomm modular acquisition developer guide 2026
- 8.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary
Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters
Anyscale
Managed Ray platform for scaling AI data processing, training, inference and RL across thousands of GPUs on any cloud
Chroma
Open-source (Apache 2.0) vector and hybrid search database for AI, with a serverless Chroma Cloud on object storage