ZML/LLMD
by ZML
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary
ZML/LLMD is a free, self-contained LLM inference server from Paris-based ZML that runs open-weight models such as Llama, Gemma, Qwen and Mistral across NVIDIA, AMD, Google TPU, Intel and Apple hardware. It is aimed at AI infrastructure teams who want to serve open models behind an OpenAI-compatible API without locking into one chip vendor.
ZML/LLMD is an LLM inference server released in alpha in July 2026 by ZML, a 20-person Paris AI-infrastructure startup founded in 2023 by Steeve Morin, formerly VP of engineering at Snap-acquired Zenly. It is built on ZML's open-source (Apache 2.0) machine-learning framework, written in Zig on top of MLIR and OpenXLA, which compiles one model codebase to multiple accelerators with no Python runtime. LLMD packages that into a self-contained server that runs Llama 2/3, Gemma 3/4, Qwen 2 through 3.6 (dense and MoE), LFM2.5 and Mistral 3/Ministral models on five architectures: NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal, with DeepSeek, Kimi, GLM, MiniMax and StepFun listed as coming soon. It ships continuous batching, paged attention, tensor-parallel sharding, prefix caching, tool calling and DFlash speculative decoding, loads weights zero-copy from Hugging Face, S3 or GCS, and exposes an OpenAI-compatible /v1/chat/completions endpoint. Distribution is via Docker images (about 1.7GB for CUDA, 280MB for TPU) and a 140MB Homebrew binary on Apple Silicon; ZML claims 1-2 second cold boots for an 8B BF16 model. ZML's own benchmarks show Gemma 4-26B at about 1,318 tok/s on two H100s and 859 tok/s on an AMD MI300X. The product is free, but unlike the framework it is not open source, and ZML has said monetisation is undecided. Alongside the launch ZML announced a $20M seed led by 20VC, with angels including Yann LeCun, Solomon Hykes and Hugging Face co-founders Clément Delangue and Julien Chaumond. It sits against vLLM, SGLang and NVIDIA's TensorRT-LLM, differentiating on hardware portability.
Platform or ML-infrastructure lead serving open-weight LLMs on a mixed or changing GPU/accelerator fleet who wants to avoid rewriting the serving stack per chip vendor.
One OpenAI-compatible inference server that runs the same open models on NVIDIA, AMD, TPU, Intel and Apple hardware, reducing hardware lock-in.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Free
- Target Market
- CTOs, ML Platform Engineers, AI Infrastructure Teams, Enterprise Developers
- Deployment
- Self-hosted, Multi-cloud, API-based
- Founded
- 2023
- Headquarters
- Paris, France
- Team Size
- 11-50
Key Features
- ✓Cross-hardware serving
Runs the same models on NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal, so capacity can be bought from any vendor.
- ✓OpenAI-compatible API
Exposes /v1/chat/completions on port 8000 as a drop-in replacement, so existing OpenAI-client code can point at self-hosted models.
- ✓Production serving techniques
Continuous batching, paged attention, tensor-parallel sharding and prefix caching raise throughput and GPU utilisation for concurrent requests.
- ✓DFlash speculative decoding
Speculative decoding on compatible models, which ZML says speeds generation by up to 10x, lowering latency per request.
- ✓Zero-copy model loading
Loads weights directly from Hugging Face, S3 or GCS without pre-downloading, simplifying deployment and speeding cold starts.
- ✓Small, Python-free footprint
Built in Zig on MLIR/OpenXLA with no Python runtime; the TPU image is about 280MB and the Mac binary about 140MB.
Capabilities
Use Cases
- •Negotiating GPU spend across vendors
An infrastructure team benchmarks the same Gemma or Qwen model on NVIDIA and AMD and buys whichever capacity is cheaper per token.
- •Self-hosted internal assistant
An enterprise serves an open model on-premises behind an OpenAI-compatible endpoint so internal apps keep data inside the network.
- •Bursty autoscaling inference
Small images and 1-2 second cold boots let a platform team scale inference replicas up and down quickly with demand.
- •Developer workstation testing
Engineers install LLMD via Homebrew on Apple Silicon Macs to test prompts and tool calling locally before deploying to GPU servers.
Ideal For
Best For
- ✓Teams evaluating AMD MI300X or Google TPU capacity as a cheaper alternative to NVIDIA for open-model serving
- ✓Self-hosting Llama, Gemma, Qwen or Mistral behind an OpenAI-compatible endpoint
- ✓Fast-scaling inference fleets that benefit from 1-2 second cold boots and small container images
- ✓Local development and testing of open models on Apple Silicon Macs via Homebrew
- ✓Multi-cloud or sovereign deployments that need the same serving stack on different accelerators
Not Ideal For
- ✗Production teams that need a mature, supported server today — LLMD is an alpha release with no published SLA or commercial support offering
- ✗Organisations that require fully open-source serving software for audit or licensing reasons; LLMD itself is proprietary even though the ZML framework is Apache 2.0
- ✗Teams running models outside the supported list (e.g. DeepSeek or Kimi, still marked coming soon)
Deployment
Market Analysis
Pros
- ✓Genuine hardware portability across NVIDIA, AMD, TPU, Intel and Apple from one server
- ✓OpenAI-compatible API makes it a drop-in for existing client code
- ✓Free, with a strong open-source framework (4.1k GitHub stars, Apache 2.0) underneath
- ✓Very small images and 1-2 second cold boots suit autoscaling
Cons
- ✗Alpha software: no published SLA, support tier or production track record, and no named customers
- ✗LLMD is proprietary even though the ZML framework is open source, and the business model is undecided
- ✗Benchmarks are vendor-published only; ZML does not publish head-to-head numbers against vLLM or SGLang, and HN commenters were still asking for DGX Spark and Strix Halo results
- ✗Model coverage is narrower than mature servers, with DeepSeek, Kimi and GLM still listed as coming soon
Pricing
LLMD (alpha)
$0
- ✓Docker images for CUDA, ROCm, TPU and oneAPI
- ✓Homebrew binary for Apple Metal
- ✓OpenAI-compatible API
- ✓Continuous batching, paged attention, prefix caching
LLMD is free to download and run on your own hardware; you pay only for the compute you run it on. There is no published paid tier, support contract or enterprise edition, and founder Steeve Morin told TechCrunch (July 2026) that monetisation is undecided. The underlying ZML framework is Apache 2.0; LLMD itself is proprietary.
Security & Compliance
Connect
Sources
This page was written from 7 sources, 5 on domains other than zml.ai.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters
Anyscale
Managed Ray platform for scaling AI data processing, training, inference and RL across thousands of GPUs on any cloud
Chroma
Open-source (Apache 2.0) vector and hybrid search database for AI, with a serverless Chroma Cloud on object storage
LiteLLM
Open-source AI gateway: 100+ LLM APIs behind one OpenAI-compatible endpoint, with cost tracking and guardrails