The best air-gapped LLM stack for most defence, health and industrial buyers is Apache-2.0 open weights — gpt-oss-120b, Qwen3 or Mistral Small — served by vLLM. Buy it through Red Hat AI if you need a support contract; run it yourself if you have a platform team. Pay NVIDIA AI Enterprise's $4,500 per GPU per year only if you are all-NVIDIA and want its pre-optimised containers. Buy Gemini on Google Distributed Cloud only for a closed frontier model with an accreditation someone else paid for. At the workload below, an 8-GPU air-gapped node costs roughly 30-40x the hosted equivalent, and even a minimal two-GPU deployment is around 10x. You air-gap because you are not allowed to call an API, never to save money.
| vLLM + Apache weights (DIY) | Red Hat AI (vLLM, supported) | NVIDIA AI Enterprise (NIM) | Gemini on GDC air-gapped | Cohere North (private) | |
|---|---|---|---|---|---|
| Verdict | Default if you have a platform team | Default if you need a vendor on the hook | Loser for most buyers — fee without a moat | Only for frontier-quality closed models | An app layer, not an inference platform |
| Software price (checked 2026-09-28) | $0 | Per physical accelerator, contact sales | $4,500/GPU/yr; $22,500/GPU perpetual | Contact sales (seat, token or per-appliance) | Contact sales |
| Models | Any weights you are licensed to run | Any vLLM-supported weights | NVIDIA-packaged NIM profiles | Gemini only | Cohere Command family |
| Offline update path | docker save, wheel mirror, weights by hash |
Mirror images to your registry | download-to-cache, copy cache |
Google-operated appliance | Vendor-delivered |
| Extra thing inside the air gap | Nothing | OpenShift, if you use OpenShift AI | A licence server (DLS) with manual uploads | Google's hardware and control plane | Cohere's stack |
| Who should NOT pick it | Teams with no one on call for GPU drivers | Shops with no OpenShift or RHEL skills | Mixed-accelerator fleets; anyone budget-bound | Anyone who wants weights they can inspect | Anyone who only needs an endpoint |
What does "air-gapped LLM" actually mean for the buyer?
An air-gapped LLM deployment is a model, its serving runtime and every dependency they touch running on a network with no route to the internet — no API calls, no weight downloads, no telemetry, no licence check-in. That last clause is where most stacks fail. The model is the easy part. The runtime, the Python packages, the tokenizer, the embedding model and the licence server all assume egress by default.
One practitioner blueprint puts it bluntly: teams going through the exercise typically discover "five to fifteen distinct egress dependencies they did not know they had" — and one it singles out is a retrieval pipeline quietly sending text to a hosted embedding API. Treat that as a practitioner's estimate, not a measured industry figure. It matches what anyone who has run tcpdump on a "sealed" GPU node will recognise.
So the comparison below is not "which model is smartest." It is: which stack leaves the fewest things inside your perimeter that want to phone home, and the cheapest path to getting a new model across the gap.
Which model weights are actually licensed for an air gap?
Pick weights under Apache 2.0, because every other licence in this category carries a clause a lawyer in a regulated buyer will want to talk about. The permissive short-list today:
- gpt-oss-120b and gpt-oss-20b — Apache 2.0, per OpenAI's repository, plus a one-line usage policy that asks only for compliance with applicable law; and the 120b model fits on "a single 80GB GPU" while the 20b runs within 16GB.
- Qwen3-235B-A22B — Apache 2.0 on its model card.
- Mistral Small 3.2 (24B) — Apache 2.0, needing about 55GB of GPU memory in bf16.
The steel-man for the others is real: Llama, Gemma and Nemotron are good models and plenty of enterprises run them. But read what you are signing.
- Llama 4 incorporates Meta's Acceptable Use Policy "by reference into this Agreement", requires you to display "Built with Llama", and sends anyone over 700 million monthly active users back to Meta for a separate licence. For most buyers the MAU cap is irrelevant; the by-reference policy is a document that can change after you have certified your system.
- Gemma's terms say Google "reserves the right to restrict (remotely or otherwise)" usage it believes violates the agreement, and also pull in a prohibited-use policy by reference. In an air gap "remotely" is moot — but your accreditor will still ask what it means.
- NVIDIA's Open Model License is commercial and royalty-free, but revocable, and your rights "automatically terminate" if you bypass or reduce the efficacy of a model's safety guardrail without a substantially similar replacement. A fine-tuning team in a defence programme needs to know that clause exists before it starts.
- Cohere's Command A open weights are CC-BY-NC — non-commercial. The download is a research artefact. Commercial use in your air gap means buying Cohere's private deployment.
The rule: the licence you can hand to an accreditor without a cover memo is Apache 2.0. Everything else is workable but costs you a legal review per model version.
How much hardware does each concurrency target need?
The memory floor is set by the model; the concurrency ceiling can only be set by load-testing your own prompts — do not trust anyone's tokens-per-second table, including this one. What you can plan from:
- Pilot, up to ~10 concurrent users. One published reference build on a Dell R740 with two 24GB Turing-era GPUs served a 4-bit Qwen3-Coder-30B-A3B at 80-120 tokens/s for one user and about 30 tokens/s per user at 10 concurrent, with time-to-first-token under 500ms warm. That is a 2018-era server. A pilot does not need Blackwell.
- Department, dozens of concurrent users on a frontier-class open model. gpt-oss-120b's single-80GB-GPU footprint means one H100/H200 per replica, and you want at least two replicas so a driver update does not take the service down. Plan for 2-4 GPUs.
- Enterprise, hundreds of concurrent users or long contexts. This is where buyers end up at an 8-GPU HGX node, which resellers put at $320,000-$420,000 for an H200 system, typically ~$370,000 in 2026 — a reseller's estimate.
Two things change the answer. Mixture-of-experts models load all their weights even when few are active, so size memory on total parameters, not active ones (we worked through that trap here). And old hardware costs engineering time: the R740 build had to force float16, avoid FP8 and run --enforce-eager because Turing lacks BF16 — each one a day someone spent reading stack traces offline.
How do you patch and update with no internet?
Every update becomes a release: download on a connected machine, verify, carry it across on approved media, re-validate, promote. The stacks differ only in how many artefacts you carry and whether the vendor has scripted it.
- vLLM DIY. The reference build's procedure is the honest template: pull everything while connected, then
docker savethe images,pip downloadthe wheels, tar the configs, and move them across on removable media with hash verification. SetHF_HUB_OFFLINE=1, which makes the Hugging Face libraries make no HTTP calls and read only cached files. Mirror your weights into a private registry rather than a shared folder (our registry comparison covers that). - Red Hat AI. OpenShift AI has a documented disconnected install: mirror the operator and serving images to your internal registry, and serve models through KServe. It is the same vLLM, packaged with a lifecycle you can put in a change ticket.
- NVIDIA AI Enterprise. NIM's air-gap guide has you run
download-to-cachefor a specific model profile on a connected system, copy the cache across, or pointNIM_MODEL_PATHat a local directory. Profiles are tuned per GPU, so a hardware change can mean a new download. And the licence itself needs infrastructure: the Delegated License Service is "fully disconnected from the NVIDIA Licensing Portal", so you download licences and upload them manually. That is one more appliance inside your perimeter and one more thing that can expire. - Gemini on GDC. Google operates the stack. That removes the update engineering and replaces it with a dependency on Google's release cadence for your model.
Budget for size. Model artefacts run 10-400GB each, and the same blueprint's warning is the one to pin on the wall: air-gapped releases "are slower by design", and teams find out when a regulator asks for the change-control evidence.
How do you evaluate and monitor with no telemetry egress?
Turn off every default phone-home, then build the eval loop entirely inside the gap — including the judge model.
Start with the defaults. vLLM collects anonymous usage data by default — hardware, model architecture, runtime settings — and it is disabled with VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1. The Hugging Face libraries send telemetry too, disabled with HF_HUB_DISABLE_TELEMETRY=1. In an air gap these fail rather than leak, but failing calls cost timeouts, and a node that is ever briefly connected for patching will send them. Set them in the base image.
Then the monitoring stack. Hosted observability does not survive isolation; the practitioner pattern is self-hosted Prometheus, Grafana and Loki, with the logs classified at the same level as the inference cluster — because the logs contain the prompts.
Evaluation is the underestimated piece. An offline eval is a fixed test set, a harness and a judge that all run inside the perimeter. Bake the benchmark data into the harness rather than pulling it at runtime, and run a smaller local judge model calibrated against your human graders. The long pole, per the same blueprint, is ground-truth annotation on on-prem tools — start that before the hardware arrives. Every model update then reruns the same suite before promotion, which is the change-control evidence your auditor wants anyway.
What does air-gapping cost versus the hosted equivalent?
Roughly 30-40x more on an enterprise node, and still around 10x on the smallest redundant deployment, before you count a single engineer — which is why cost cannot be the reason you do it.
The defined workload: gpt-oss-120b, 1 billion tokens a month (700M input, 300M output), on one 8-GPU H200 node. That node is sized for peak concurrency and headroom, not the average: 1 billion tokens a month is under 400 tokens a second on average, which two GPUs can serve. Right-sized to two GPUs — roughly a quarter of the node, ~$31,000 a year — the gap narrows to about 9x. It does not close.
- Hosted. Together AI lists gpt-oss-120b at $0.15 per million input tokens and $0.60 per million output tokens (checked 2026-09-28). That is $105 + $180 = $285 a month, about $3,400 a year.
- Air-gapped hardware. A ~$370,000 H200 node straight-lined over three years is ~$123,000 a year, before power, cooling, spares and rack space.
- Air-gapped software. Free for DIY vLLM. Red Hat is per physical accelerator and unpriced publicly — get a quote. NVIDIA AI Enterprise lists $4,500 per GPU for one year, $18,000 for five, or $22,500 perpetual with five years of support — $36,000 a year for this node.
Look at that last line against the hosted bill. At ten times the volume — 10 billion tokens a month — hosted gpt-oss-120b costs about $34,200 a year. The NVIDIA licence alone, for one node, costs more than the entire hosted service at 10x the workload. And that is before the people: GPU driver lifecycles, offline release engineering, the eval pipeline and the on-call rotation.
The working title of this piece was that air-gapped LLMs cost more in operations than they save in licences. The numbers say it is worse: there are no licence savings to set against the operations bill, because the hosted open-weight API was never expensive. The only lever you control is the recurring software line — which is why the verdict pushes that line to zero or near it.
Where each option wins, and who should not buy it
vLLM with Apache weights, self-run. The default for any team that already runs GPUs. vLLM is also our default inference runtime outside air gaps, so you are not maintaining an exotic stack. Do not pick it if nobody on staff will own GPU drivers, CUDA versions and a pip mirror at 02:00 with no internet to search.
Red Hat AI. The same vLLM with a support contract, a documented disconnected installer and a subscription counted per physical accelerator — eight GPUs, eight subscriptions; OpenShift AI additionally needs base OpenShift subscriptions. For a hospital or a defence integrator whose procurement rules require a named vendor, this is the pick. Do not pick it if you have no OpenShift or RHEL skills — you will be buying a platform to host an endpoint.
NVIDIA AI Enterprise (NIM). The steel-man: pre-optimised containers per GPU, one throat to choke on an all-NVIDIA estate, and a scripted air-gap procedure. This is the loser for most buyers. It charges the largest published recurring fee in the comparison for packaging open runtimes and open weights you can mirror yourself, and it adds a licence server with manual uploads to the one environment where every extra component is a liability. Do not pick it if you run mixed accelerators, if your budget is fixed, or if your accreditor counts appliances.
Gemini on Google Distributed Cloud air-gapped. Google made Gemini on GDC air-gapped generally available in August 2025, says the product is authorised for US Government Secret and Top Secret missions, and at Next '26 partner Cirrascale announced a single-server Dell appliance with eight NVIDIA GPUs, priced by seat, by token or flat per appliance — contact sales. It is the only way to get a closed frontier model behind a gap. Do not pick it if you need weights you can inspect, fine-tune or keep after the contract ends; the model leaves when the appliance does.
Cohere North. Cohere offers air-gapped deployment behind your firewall, and North was launched as runnable on as few as two GPUs. It is an agent-and-search application with models attached, priced through sales. Do not pick it if what you need is an OpenAI-compatible endpoint for your own applications — you would be buying an application to get a model whose open weights you cannot use commercially.
The decision: criteria that predict regret
The buyers who regret an air-gapped LLM purchase almost always under-weighted operations and over-weighted the model. Score these before you sign:
- Can you name the person who carries the next model across the gap? If not, buy Red Hat or Google, not DIY.
- Does the licence survive your accreditor without a memo? Apache 2.0 does. Anything with a policy incorporated by reference costs a legal review per version.
- How many components phone home by default? Count them in a connected staging environment with egress logging on. Every one you find there is one you will not find in production.
- What is your recurring software line per GPU? At air-gap volumes it is the biggest cost you can still change.
- Do you need a closed frontier model, or just a good one? If an open 120B model passes your offline eval, the Gemini appliance is a premium for a gap you have not measured.
What changes the answer: a mandate for a closed frontier model (Gemini on GDC); a procurement rule requiring a single hardware-and-software vendor on an all-NVIDIA estate (NIM becomes defensible); or a use case that is really search-and-agents over internal documents with no engineering team (North).
This Week: stand up gpt-oss-20b on one 16GB GPU in a staging enclave with egress logging on, and write down every blocked connection.
This Month: build the offline eval suite — 200 real prompts, graded by your people — and run gpt-oss-120b, Qwen3 and one paid option against it.
Before Budget Close: get a Red Hat per-accelerator quote and an NVIDIA AI Enterprise quote for the same node, and put both next to the $0 line.
The Bottom Line
Air-gapped AI repeats the pattern of every sovereign-IT cycle before it: the hardware is a one-time decision, and the software subscription and the release process are the cost that compounds. The hosted open-weight API is cheap enough that nothing inside your perimeter will ever beat it on price — so stop trying, and make the perimeter as cheap to operate as possible. Apache weights, vLLM, a support contract only if your rules require one.
The gap is the requirement. The fee is optional.
Continue Reading
- Hugging Face Alternatives: Cache the Hub, Sign at the Door
- vLLM vs TensorRT-LLM vs SGLang: Default to vLLM
- Tencent Says 49B Active. Your Node Loads All 770B.
- 7 OpenAI Alternatives. Only 3 Clear a Sovereignty Rule.
- Best MLOps for Regulated AI: Domino, Then a Sign-Off Layer
- Self-Host the Vector DB for Residency. Not for the Bill.
