Z

ZML/LLMD

by ZML

Infrastructure & CloudDeveloper ToolsAI Models & APIs

Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary

Free·Added Sep 28, 2026·Updated Sep 28, 2026
Share:
THE DAILY BRIEF
ZML/LLMD

by ZML

Infrastructure & CloudDeveloper ToolsAI Models & APIs

Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary

Free

ZML/LLMD is a free, self-contained LLM inference server from Paris-based ZML that runs open-weight models such as Llama, Gemma, Qwen and Mistral across NVIDIA, AMD, Google TPU, Intel and Apple hardware. It is aimed at AI infrastructure teams who want to serve open models behind an OpenAI-compatible API without locking into one chip vendor.

At a Glance

Category
Infrastructure & Cloud
Pricing
Free
Target Market
CTOs, ML Platform Engineers, AI Infrastructure Teams, Enterprise Developers
Deployment
Self-hosted, Multi-cloud, API-based
Founded
2023
Headquarters
Paris, France
Team Size
11-50

Key Features

  • ✓Cross-hardware serving
  • ✓OpenAI-compatible API
  • ✓Production serving techniques
  • ✓DFlash speculative decoding
  • ✓Zero-copy model loading
  • ✓Small, Python-free footprint

Capabilities

✓text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •Negotiating GPU spend across vendors
  • •Self-hosted internal assistant
  • •Bursty autoscaling inference
  • •Developer workstation testing

Ideal For

Best For

  • ✓Teams evaluating AMD MI300X or Google TPU capacity as a cheaper alternative to NVIDIA for open-model serving
  • ✓Self-hosting Llama, Gemma, Qwen or Mistral behind an OpenAI-compatible endpoint
  • ✓Fast-scaling inference fleets that benefit from 1-2 second cold boots and small container images
  • ✓Local development and testing of open models on Apple Silicon Macs via Homebrew
  • ✓Multi-cloud or sovereign deployments that need the same serving stack on different accelerators

Not Ideal For

  • ✗Production teams that need a mature, supported server today — LLMD is an alpha release with no published SLA or commercial support offering
  • ✗Organisations that require fully open-source serving software for audit or licensing reasons; LLMD itself is proprietary even though the ZML framework is Apache 2.0
  • ✗Teams running models outside the supported list (e.g. DeepSeek or Kimi, still marked coming soon)

Market Analysis

Open-model inferenceHardware-agnosticDeveloper-first

Pros

  • ✓Genuine hardware portability across NVIDIA, AMD, TPU, Intel and Apple from one server
  • ✓OpenAI-compatible API makes it a drop-in for existing client code
  • ✓Free, with a strong open-source framework (4.1k GitHub stars, Apache 2.0) underneath
  • ✓Very small images and 1-2 second cold boots suit autoscaling

Cons

  • ✗Alpha software: no published SLA, support tier or production track record, and no named customers
  • ✗LLMD is proprietary even though the ZML framework is open source, and the business model is undecided
  • ✗Benchmarks are vendor-published only; ZML does not publish head-to-head numbers against vLLM or SGLang, and HN commenters were still asking for DGX Spark and Strix Halo results
  • ✗Model coverage is narrower than mature servers, with DeepSeek, Kimi and GLM still listed as coming soon

Pricing

LLMD (alpha)

$0

  • ✓Docker images for CUDA, ROCm, TPU and oneAPI
  • ✓Homebrew binary for Apple Metal
  • ✓OpenAI-compatible API
  • ✓Continuous batching, paged attention, prefix caching

LLMD is free to download and run on your own hardware; you pay only for the compute you run it on. There is no published paid tier, support contract or enterprise edition, and founder Steeve Morin told TechCrunch (July 2026) that monetisation is undecided. The underlying ZML framework is Apache 2.0; LLMD itself is proprietary.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

ZML/LLMD is a free, self-contained LLM inference server from Paris-based ZML that runs open-weight models such as Llama, Gemma, Qwen and Mistral across NVIDIA, AMD, Google TPU, Intel and Apple hardware. It is aimed at AI infrastructure teams who want to serve open models behind an OpenAI-compatible API without locking into one chip vendor.

ZML/LLMD is an LLM inference server released in alpha in July 2026 by ZML, a 20-person Paris AI-infrastructure startup founded in 2023 by Steeve Morin, formerly VP of engineering at Snap-acquired Zenly. It is built on ZML's open-source (Apache 2.0) machine-learning framework, written in Zig on top of MLIR and OpenXLA, which compiles one model codebase to multiple accelerators with no Python runtime. LLMD packages that into a self-contained server that runs Llama 2/3, Gemma 3/4, Qwen 2 through 3.6 (dense and MoE), LFM2.5 and Mistral 3/Ministral models on five architectures: NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal, with DeepSeek, Kimi, GLM, MiniMax and StepFun listed as coming soon. It ships continuous batching, paged attention, tensor-parallel sharding, prefix caching, tool calling and DFlash speculative decoding, loads weights zero-copy from Hugging Face, S3 or GCS, and exposes an OpenAI-compatible /v1/chat/completions endpoint. Distribution is via Docker images (about 1.7GB for CUDA, 280MB for TPU) and a 140MB Homebrew binary on Apple Silicon; ZML claims 1-2 second cold boots for an 8B BF16 model. ZML's own benchmarks show Gemma 4-26B at about 1,318 tok/s on two H100s and 859 tok/s on an AMD MI300X. The product is free, but unlike the framework it is not open source, and ZML has said monetisation is undecided. Alongside the launch ZML announced a $20M seed led by 20VC, with angels including Yann LeCun, Solomon Hykes and Hugging Face co-founders Clément Delangue and Julien Chaumond. It sits against vLLM, SGLang and NVIDIA's TensorRT-LLM, differentiating on hardware portability.

Ideal Buyer

Platform or ML-infrastructure lead serving open-weight LLMs on a mixed or changing GPU/accelerator fleet who wants to avoid rewriting the serving stack per chip vendor.

Key Benefit

One OpenAI-compatible inference server that runs the same open models on NVIDIA, AMD, TPU, Intel and Apple hardware, reducing hardware lock-in.

At a Glance

Category
Infrastructure & Cloud
Pricing
Free
Target Market
CTOs, ML Platform Engineers, AI Infrastructure Teams, Enterprise Developers
Deployment
Self-hosted, Multi-cloud, API-based
Founded
2023
Headquarters
Paris, France
Team Size
11-50

Key Features

  • ✓
    Cross-hardware serving

    Runs the same models on NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal, so capacity can be bought from any vendor.

  • ✓
    OpenAI-compatible API

    Exposes /v1/chat/completions on port 8000 as a drop-in replacement, so existing OpenAI-client code can point at self-hosted models.

  • ✓
    Production serving techniques

    Continuous batching, paged attention, tensor-parallel sharding and prefix caching raise throughput and GPU utilisation for concurrent requests.

  • ✓
    DFlash speculative decoding

    Speculative decoding on compatible models, which ZML says speeds generation by up to 10x, lowering latency per request.

  • ✓
    Zero-copy model loading

    Loads weights directly from Hugging Face, S3 or GCS without pre-downloading, simplifying deployment and speeding cold starts.

  • ✓
    Small, Python-free footprint

    Built in Zig on MLIR/OpenXLA with no Python runtime; the TPU image is about 280MB and the Mac binary about 140MB.

Capabilities

✓text generation
✗image generation
✗video generation
✗code generation
✗workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •
    Negotiating GPU spend across vendors

    An infrastructure team benchmarks the same Gemma or Qwen model on NVIDIA and AMD and buys whichever capacity is cheaper per token.

  • •
    Self-hosted internal assistant

    An enterprise serves an open model on-premises behind an OpenAI-compatible endpoint so internal apps keep data inside the network.

  • •
    Bursty autoscaling inference

    Small images and 1-2 second cold boots let a platform team scale inference replicas up and down quickly with demand.

  • •
    Developer workstation testing

    Engineers install LLMD via Homebrew on Apple Silicon Macs to test prompts and tool calling locally before deploying to GPU servers.

Ideal For

Best For

  • ✓Teams evaluating AMD MI300X or Google TPU capacity as a cheaper alternative to NVIDIA for open-model serving
  • ✓Self-hosting Llama, Gemma, Qwen or Mistral behind an OpenAI-compatible endpoint
  • ✓Fast-scaling inference fleets that benefit from 1-2 second cold boots and small container images
  • ✓Local development and testing of open models on Apple Silicon Macs via Homebrew
  • ✓Multi-cloud or sovereign deployments that need the same serving stack on different accelerators

Not Ideal For

  • ✗Production teams that need a mature, supported server today — LLMD is an alpha release with no published SLA or commercial support offering
  • ✗Organisations that require fully open-source serving software for audit or licensing reasons; LLMD itself is proprietary even though the ZML framework is Apache 2.0
  • ✗Teams running models outside the supported list (e.g. DeepSeek or Kimi, still marked coming soon)

Deployment

✓On-Premise

Market Analysis

Open-model inferenceHardware-agnosticDeveloper-first

Pros

  • ✓Genuine hardware portability across NVIDIA, AMD, TPU, Intel and Apple from one server
  • ✓OpenAI-compatible API makes it a drop-in for existing client code
  • ✓Free, with a strong open-source framework (4.1k GitHub stars, Apache 2.0) underneath
  • ✓Very small images and 1-2 second cold boots suit autoscaling

Cons

  • ✗Alpha software: no published SLA, support tier or production track record, and no named customers
  • ✗LLMD is proprietary even though the ZML framework is open source, and the business model is undecided
  • ✗Benchmarks are vendor-published only; ZML does not publish head-to-head numbers against vLLM or SGLang, and HN commenters were still asking for DGX Spark and Strix Halo results
  • ✗Model coverage is narrower than mature servers, with DeepSeek, Kimi and GLM still listed as coming soon

Pricing

LLMD (alpha)

$0

  • ✓Docker images for CUDA, ROCm, TPU and oneAPI
  • ✓Homebrew binary for Apple Metal
  • ✓OpenAI-compatible API
  • ✓Continuous batching, paged attention, prefix caching

LLMD is free to download and run on your own hardware; you pay only for the compute you run it on. There is no published paid tier, support contract or enterprise edition, and founder Steeve Morin told TechCrunch (July 2026) that monetisation is undecided. The underlying ZML framework is Apache 2.0; LLMD itself is proprietary.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

Connect

Sources

This page was written from 7 sources, 5 on domains other than zml.ai.

  1. 1.zml.ai — zml.aivendor
  2. 2.zml.ai — llmdvendor
  3. 3.techcrunch.com — hot french startup zml releases free product to speed infere
  4. 4.frenchtechjournal.com — french tech funding wire july 13 skello eur150m led 12 deals
  5. 5.github.com — zml
  6. 6.hub.docker.com — llmd
  7. 7.news.ycombinator.com — item
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe