ModelsFrameworks

AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference

by vLLM Project (AIBrix Team)

AdvancedDocumentationFreeSelf-paced reference; ~1-2 hours for the quickstart on a GPU cluster, several days to cover routing, autoscaling and KV cache offloading

Run vLLM in production on Kubernetes with LLM-aware routing, prefill-decode disaggregation, LoRA packing, autoscaling and distributed KV cache.

Start LearningAdded Oct 8, 2026 · Updated Oct 8, 2026

Overview

AIBrix is an Apache 2.0 project hosted under the vllm-project GitHub organization that supplies cloud-native building blocks for large-scale LLM inference, and its Read the Docs site is the canonical guide. The documentation is split into getting started (concepts, a quickstart, container images, installation, advanced Kubernetes examples, FAQ), architecture (router, engine runtime, autoscaler, KV cache offloading framework, StormService), gateway and routing (gateway routing, external replica routing, prefill-decode disaggregation, agentic routing, the vLLM Semantic Router, vLLM-Omni multimodal serving), model serving (dynamic LoRA loading, multi-node inference, multi-engine support, experimental heterogeneous GPU inference, ModelClaim GPU runtime pools, model warmup), scaling and performance, batch inference with a Batch API, the Brixbench benchmark and workload generator, and production readiness (gateway deployment, observability, the AIBrix Console). The v0.7.0 quickstart installs AIBrix with three kubectl manifests or a Helm chart, deploys DeepSeek-R1-Distill-Llama-8B on vLLM, and queries an Envoy gateway through OpenAI-compatible endpoints, choosing a routing strategy (random, least-request, prefix-cache or pd) with a request header. The design is described in a February 2025 arXiv paper by the AIBrix Team, which reports that its distributed KV cache raised throughput by 50% and cut inference latency by 70% in the authors' tests. The repo has about 5.1k stars, and v0.7.0 shipped on 2026-06-16 after v0.6.0 in March 2026.

At a Glance

Topic
Models
Level
Advanced
Format
Documentation
Cost
Free
Duration
Self-paced reference; ~1-2 hours for the quickstart on a GPU cluster, several days to cover routing, autoscaling and KV cache offloading
Provider
vLLM Project (AIBrix Team)
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Install the AIBrix control plane on Kubernetes with kubectl manifests or the Helm chart
  • ✓Deploy a vLLM model behind the AIBrix Envoy gateway and query it with OpenAI-compatible APIs
  • ✓Choose between random, least-request, prefix-cache and prefill-decode routing strategies per request
  • ✓Run prefill-decode disaggregation with StormService roles and NIXL-enabled vLLM images
  • ✓Load many LoRA adapters dynamically onto shared base models for high-density serving
  • ✓Configure LLM-specific autoscaling and KV cache offloading to cut latency and GPU cost
  • ✓Serve large models across multiple nodes and mix engines or heterogeneous GPU types
  • ✓Benchmark inference throughput and resource efficiency with Brixbench and the workload generator

Highlights

  • •Lives under the official vllm-project GitHub organization and is co-designed with the vLLM engine
  • •Covers production problems most serving docs skip: prefix-cache-aware routing, PD disaggregation, LoRA density and GPU failure detection
  • •Backed by an arXiv systems paper reporting a 50% throughput gain and 70% latency reduction from its distributed KV cache
  • •Steady release cadence with dated notes for v0.4.0 through v0.7.0 (June 2026) and about 5.1k GitHub stars
  • •Includes a batch inference API and a benchmarking toolkit, so capacity tests use the same stack as production

Who It’s For

Best For

  • ✓Platform and MLOps engineers running vLLM on shared Kubernetes GPU clusters
  • ✓Teams serving many fine-tuned LoRA variants that need to pack them onto fewer GPUs
  • ✓Infrastructure engineers evaluating prefill-decode disaggregation and KV cache reuse to lower cost per token
  • ✓Engineers comparing Kubernetes-native LLM serving stacks such as llm-d, KServe and Dynamo

Prerequisites

  • •Hands-on Kubernetes experience: kubectl, CRDs, Services and Helm
  • •Access to a Kubernetes cluster with NVIDIA GPUs and the device plugin installed
  • •Familiarity with serving models on vLLM and with LLM inference basics such as KV cache and batching

FAQ

What is AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference?

The official documentation for AIBrix, an open-source control plane from the vLLM project for deploying, routing and scaling LLM inference on Kubernetes. It is for platform and MLOps engineers who already serve models with vLLM and now need a gateway, autoscaling, multi-node serving and KV cache reuse across a shared GPU fleet.

Is AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference free?

AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference is free to access.

What level is AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference for?

AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference is aimed at a advanced audience. Recommended background: Hands-on Kubernetes experience: kubectl, CRDs, Services and Helm, Access to a Kubernetes cluster with NVIDIA GPUs and the device plugin installed, Familiarity with serving models on vLLM and with LLM inference basics such as KV cache and batching.

How long does AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference take?

Expect roughly Self-paced reference; ~1-2 hours for the quickstart on a GPU cluster, several days to cover routing, autoscaling and KV cache offloading. Most learners work through it at their own pace.

What will I learn from AIBrix Documentation: Kubernetes-Native Building Blocks for Scalable LLM Inference?

You'll learn: Install the AIBrix control plane on Kubernetes with kubectl manifests or the Helm chart; Deploy a vLLM model behind the AIBrix Envoy gateway and query it with OpenAI-compatible APIs; Choose between random, least-request, prefix-cache and prefill-decode routing strategies per request; Run prefill-decode disaggregation with StormService roles and NIXL-enabled vLLM images; Load many LoRA adapters dynamically onto shared base models for high-density serving; Configure LLM-specific autoscaling and KV cache offloading to cut latency and GPU cost; Serve large models across multiple nodes and mix engines or heterogeneous GPU types; Benchmark inference throughput and resource efficiency with Brixbench and the workload generator.

Topics

aibrixvllmllm-inferencekuberneteskv-cachelora

Sources

This page was written from 4 sources, 2 on domains other than aibrix.readthedocs.io.

  1. 1.aibrix.readthedocs.io — latestvendor
  2. 2.aibrix.readthedocs.io — quickstartvendor
  3. 3.github.com — aibrix
  4. 4.arxiv.org — 2504.03648