M

Modular

by Modular (a Qualcomm company)

Infrastructure & CloudDeveloper ToolsAI Models & APIs

MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips

Free · Usage-based · Contact for pricing·Added Oct 2, 2026·Updated Oct 2, 2026
Share:
THE DAILY BRIEF
Modular

by Modular (a Qualcomm company)

Infrastructure & CloudDeveloper ToolsAI Models & APIs

MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips

Free · Usage-based · Contact for pricing

Modular builds MAX, an inference serving framework, and Mojo, a Python-like systems language for GPU kernels, plus a managed Modular Cloud. It is for platform and ML infrastructure teams that want to serve open models on several hardware vendors from one stack instead of being tied to CUDA.

At a Glance

Category
Infrastructure & Cloud
Pricing
Free, Usage-based, Contact for pricing
Target Market
CTOs, ML Platform Teams, AI Infrastructure Engineers, Enterprise Developers
Deployment
Open-source, Self-hosted, Multi-cloud, Hybrid
Founded
2022
Headquarters
Palo Alto, United States
Team Size
51-200

Key Features

  • ✓MAX inference server
  • ✓Mojo 1.0 language
  • ✓Open-source compiler
  • ✓Multi-vendor hardware support
  • ✓Modular Cloud endpoints
  • ✓Bring Your Own Cloud
  • ✓Forward-deployed engineers

Capabilities

✓text generation
✓image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •Cutting inference cost on AMD
  • •Low-latency voice AI
  • •Writing portable custom kernels
  • •Private inference in your VPC

Ideal For

Best For

  • ✓Serving open-weight LLMs on AMD GPUs alongside or instead of NVIDIA
  • ✓Teams writing custom GPU kernels who want one portable language instead of CUDA and ROCm
  • ✓Running inference inside your own VPC with vendor engineers tuning the deployment
  • ✓Low-latency voice and chat workloads that need sub-second time to first token
  • ✓Free self-hosted inference with OpenAI-compatible endpoints

Not Ideal For

  • ✗Teams that need hardware neutrality guaranteed for years, since Qualcomm now owns the stack and analysts flag a risk that its own accelerators get optimisation priority
  • ✗Python shops expecting Mojo to run existing Python code; Hacker News practitioners note source compatibility with Python is not close
  • ✗Buyers who need published enterprise pricing; BYOC and Enterprise are quote-only

Market Analysis

Hardware-agnostic inferenceOpen-source

Pros

  • ✓Genuine multi-vendor GPU support reduces dependence on NVIDIA supply and pricing
  • ✓Open-source compiler and kernels with an active community of roughly 200 contributors on the 1.0 release
  • ✓Free self-hosted edition with OpenAI-compatible endpoints lowers the cost of a trial
  • ✓Named production users including Inworld, MiniMax and Hippocratic AI

Cons

  • ✗Qualcomm ownership raises questions about long-term hardware neutrality and future pricing
  • ✗Mojo looks like Python but does not run Python source, and some Hacker News users said they never got past toy examples
  • ✗Performance claims against vLLM and CUDA are vendor-reported and lack independent verification
  • ✗Licensing scope for MAX outside the open-sourced components was still unclear after the August 2026 release, per RuntimeWire

Pricing

Self-Hosted

$0

  • ✓MAX and Mojo under community license
  • ✓Deploy anywhere
  • ✓Community support via Discord and GitHub

Modular Cloud

Usage-based (e.g. Qwen 3.5 9B $0.17/$0.25 per 1M tokens)

  • ✓Shared endpoints billed per token
  • ✓Dedicated endpoints billed per GPU minute
  • ✓SOC 2 Type 2
  • ✓Forward-deployed engineers

Bring Your Own Cloud

Contact for pricing

  • ✓Billed per minute of reserved GPU capacity
  • ✓Runs in your VPC
  • ✓Custom APIs

Enterprise

Contact for pricing

  • ✓AWS, GCP, Azure or Oracle
  • ✓Hybrid deployment
  • ✓Custom engagement

Self-hosting MAX and Mojo is free. Modular Cloud publishes per-token rates for shared endpoints (DeepSeek V4 at $1.74 input / $3.48 output per 1M tokens, Qwen 3.5 9B at $0.17 / $0.25) and bills dedicated endpoints per GPU minute. BYOC and Enterprise are quote-only, and one analysis warns enterprise pricing could change after the Qualcomm integration. Prices checked on modular.com/pricing on October 2, 2026.

Security & Compliance

✓soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✓data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Modular builds MAX, an inference serving framework, and Mojo, a Python-like systems language for GPU kernels, plus a managed Modular Cloud. It is for platform and ML infrastructure teams that want to serve open models on several hardware vendors from one stack instead of being tied to CUDA.

Modular was founded in 2022 by Chris Lattner, the creator of LLVM and Swift, and Tim Davis, and is based in Palo Alto. Its stack has three parts. MAX is a model serving and modeling framework with OpenAI-compatible endpoints that ships as a container under 700 MB and runs on NVIDIA and AMD GPUs and x86 and ARM CPUs; the company also lists TPU, Trainium, Intel, Qualcomm and Apple silicon support. Mojo is a Python-styled systems language for writing portable GPU kernels; Mojo 1.0 shipped on August 11, 2026 in the Modular 26.5 release, and the compiler was open-sourced under Apache 2.0 with LLVM exceptions on August 18. Modular Cloud runs the same stack as shared per-token endpoints, dedicated per-GPU-minute endpoints, or inside a customer's own VPC. Modular raised a $250 million Series C at a $1.6 billion valuation in September 2025, led by Thomas Tull's USIT fund, bringing total funding to $380 million. Qualcomm completed its acquisition on July 29, 2026 in an all-stock deal reported at $3.9 billion, and says Mojo, MAX and Modular Cloud continue as products and that MAX will keep supporting third-party hardware. Named users include Inworld, MiniMax, Hippocratic AI, TensorWave and AWS. The vendor claims up to 2x the throughput of vLLM on the same hardware; those figures have not been independently verified. The main GitHub repository has about 29,900 stars and 3,200 forks.

Ideal Buyer

A head of ML platform running open-weight models at scale who wants to buy AMD or other non-NVIDIA accelerators without rewriting the serving stack.

Key Benefit

One inference stack and one kernel language that runs across GPU vendors, with self-hosting free and a managed cloud for burst capacity.

At a Glance

Category
Infrastructure & Cloud
Pricing
Free, Usage-based, Contact for pricing
Target Market
CTOs, ML Platform Teams, AI Infrastructure Engineers, Enterprise Developers
Deployment
Open-source, Self-hosted, Multi-cloud, Hybrid
Founded
2022
Headquarters
Palo Alto, United States
Team Size
51-200

Key Features

  • ✓
    MAX inference server

    Serves 1,000+ pre-configured models behind OpenAI-compatible REST endpoints from a container under 700 MB.

  • ✓
    Mojo 1.0 language

    A Python-styled systems language for writing GPU kernels once and running them across chip vendors, now stable at 1.0.

  • ✓
    Open-source compiler

    The Mojo compiler, standard library and MAX kernels are on GitHub under Apache 2.0 with LLVM exceptions, so teams can audit and fork them.

  • ✓
    Multi-vendor hardware support

    Runs on NVIDIA and AMD GPUs and x86 and ARM CPUs, with TPU, Trainium, Intel, Qualcomm and Apple silicon also listed, which reduces lock-in to one accelerator.

  • ✓
    Modular Cloud endpoints

    Shared per-token endpoints for testing and dedicated per-GPU-minute endpoints for production, with SOC 2 Type 2 certification.

  • ✓
    Bring Your Own Cloud

    Deploys the stack in the customer's VPC on their cloud credits, so data never leaves their environment.

  • ✓
    Forward-deployed engineers

    Cloud and BYOC plans include Modular engineers who tune throughput and latency for the customer's specific pipelines.

Capabilities

✓text generation
✓image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •
    Cutting inference cost on AMD

    Move open-model serving onto AMD GPUs without rewriting kernels; TensorWave reports 70% total cost savings on AMD with MAX.

  • •
    Low-latency voice AI

    Serve speech and LLM models with sub-second time to first token, as Inworld and Hippocratic AI do for real-time conversations.

  • •
    Writing portable custom kernels

    Write performance-critical GPU kernels in Mojo once and run them on several vendors instead of maintaining CUDA and ROCm versions.

  • •
    Private inference in your VPC

    Run the managed stack inside your own AWS, GCP, Azure or Oracle account to satisfy data residency and compliance policies.

Ideal For

Best For

  • ✓Serving open-weight LLMs on AMD GPUs alongside or instead of NVIDIA
  • ✓Teams writing custom GPU kernels who want one portable language instead of CUDA and ROCm
  • ✓Running inference inside your own VPC with vendor engineers tuning the deployment
  • ✓Low-latency voice and chat workloads that need sub-second time to first token
  • ✓Free self-hosted inference with OpenAI-compatible endpoints

Not Ideal For

  • ✗Teams that need hardware neutrality guaranteed for years, since Qualcomm now owns the stack and analysts flag a risk that its own accelerators get optimisation priority
  • ✗Python shops expecting Mojo to run existing Python code; Hacker News practitioners note source compatibility with Python is not close
  • ✗Buyers who need published enterprise pricing; BYOC and Enterprise are quote-only

Integrations

✓SDK Available
SDK:PythonMojo

Deployment

✓On-Premise

Market Analysis

Hardware-agnostic inferenceOpen-source

Pros

  • ✓Genuine multi-vendor GPU support reduces dependence on NVIDIA supply and pricing
  • ✓Open-source compiler and kernels with an active community of roughly 200 contributors on the 1.0 release
  • ✓Free self-hosted edition with OpenAI-compatible endpoints lowers the cost of a trial
  • ✓Named production users including Inworld, MiniMax and Hippocratic AI

Cons

  • ✗Qualcomm ownership raises questions about long-term hardware neutrality and future pricing
  • ✗Mojo looks like Python but does not run Python source, and some Hacker News users said they never got past toy examples
  • ✗Performance claims against vLLM and CUDA are vendor-reported and lack independent verification
  • ✗Licensing scope for MAX outside the open-sourced components was still unclear after the August 2026 release, per RuntimeWire

Pricing

✓Free Trial Available

Self-Hosted

$0

  • ✓MAX and Mojo under community license
  • ✓Deploy anywhere
  • ✓Community support via Discord and GitHub

Modular Cloud

Usage-based (e.g. Qwen 3.5 9B $0.17/$0.25 per 1M tokens)

  • ✓Shared endpoints billed per token
  • ✓Dedicated endpoints billed per GPU minute
  • ✓SOC 2 Type 2
  • ✓Forward-deployed engineers

Bring Your Own Cloud

Contact for pricing

  • ✓Billed per minute of reserved GPU capacity
  • ✓Runs in your VPC
  • ✓Custom APIs

Enterprise

Contact for pricing

  • ✓AWS, GCP, Azure or Oracle
  • ✓Hybrid deployment
  • ✓Custom engagement

Self-hosting MAX and Mojo is free. Modular Cloud publishes per-token rates for shared endpoints (DeepSeek V4 at $1.74 input / $3.48 output per 1M tokens, Qwen 3.5 9B at $0.17 / $0.25) and bills dedicated endpoints per GPU minute. BYOC and Enterprise are quote-only, and one analysis warns enterprise pricing could change after the Qualcomm integration. Prices checked on modular.com/pricing on October 2, 2026.

Security & Compliance

✓soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✓data residency

Connect

Sources

This page was written from 8 sources, 5 on domains other than modular.com.

  1. 1.modular.com — modular.comvendor
  2. 2.modular.com — pricingvendor
  3. 3.modular.com — modular 26 5 mojo 1 0 is herevendor
  4. 4.github.com — modular
  5. 5.sdxcentral.com — modular raises 250m for ais unified compute layer at 16b val
  6. 6.runtimewire.com — chris lattner open sources mojo qualcomm modular
  7. 7.developersdigest.tech — qualcomm modular acquisition developer guide 2026
  8. 8.hn.algolia.com — search
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe