ModelsAgentic

jamesob's Guide to Running SOTA LLMs Locally

by James O'Beirne

AdvancedGuideFree~1-2 hour read; the full multi-GPU build is a multi-week hardware project

A real build log for serving frontier open-weight LLMs at home: budget tiers, 4x RTX PRO 6000, PCIe switches, vLLM and the BIOS/kernel fixes that make P2P work.

Start LearningAdded Sep 30, 2026 · Updated Sep 30, 2026

Overview

jamesob's guide to running SOTA LLMs locally is a GitHub README-plus-configs written by James O'Beirne in mid-2026 that reached 412 points and 181 comments on Hacker News in July 2026 and has roughly 1,900 stars. It is structured as a sequence of sections: how much to spend (a ~$2k tier of two RTX 3090s with 48GB VRAM running Qwen3.6-27B plus cohere-transcribe for speech-to-text, a still-TODO ~$20k tier, and a ~$40k+ tier), the base system (ASRock Rack ROMED8-2T, AMD EPYC Milan 7313P, 128GB DDR4 ECC, dual 1700W PSUs, about $5.7k sourced from eBay), four 96GB RTX PRO 6000 Blackwell GPUs for 384GB VRAM, a c-payne Microchip Switchtec PCIe Gen4 switch sub-BOM, hoarding weights with hf download onto mirrored ZFS, and running each model in its own Docker container with vLLM, including a GLM-5.2-Int8Mix-NVFP4-REAP-594B config at about 80 tokens/s with 460k context. It then documents the hard-won fixes: forcing Gen4 and disabling ASPM in BIOS, iommu=off because NCCL otherwise hangs on multi-GPU P2P, a systemd oneshot that disables ACS so nvidia-smi topo shows PIX between GPUs, and capping each GPU at 350W to fit a 110V circuit, ending with measured 27.5 GB/s unidirectional P2P bandwidth. A final section describes the agent harness on top: opencode in tmux, web search via SearXNG and Kagi, a Telegram bot, local Gitea and a sandboxed VM.

At a Glance

Topic
Models
Level
Advanced
Format
Guide
Cost
Free
Duration
~1-2 hour read; the full multi-GPU build is a multi-week hardware project
Provider
James O'Beirne
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Match a hardware budget tier to the open-weight models it can realistically serve
  • ✓Spec a multi-GPU EPYC workstation with RTX PRO 6000 cards and adequate power
  • ✓Use an indie PCIe Gen4 switch to get fast GPU peer-to-peer bandwidth
  • ✓Configure BIOS, GRUB and ACS settings so NCCL multi-GPU P2P stops hanging
  • ✓Run each model in its own vLLM Docker container with read-only weight mounts
  • ✓Serve heavily quantized and REAP-pruned frontier MoE models with very long contexts
  • ✓Power-limit GPUs with nvidia-smi to fit a household electrical circuit safely
  • ✓Wire a self-hosted coding-agent harness with local search, Git hosting and sandboxing

Highlights

  • •A real, reproducible bill of materials with prices, not a generic 'best GPU for LLMs' listicle
  • •Documents the undocumented failures (NCCL P2P hangs, ASPM link downgrades, ACS) and their exact fixes
  • •Includes working docker-compose vLLM runners and a GPU P2P benchmark script in the repo
  • •The 181-comment Hacker News thread is an essential companion: practitioners flag that heavily quantized/pruned models degrade on long-context tasks and debate whether $40k+ beats paying for API tokens
  • •Covers a cheap ~$2k two-3090 tier as well as the high-end build

Who It’s For

Best For

  • ✓ML infrastructure engineers building on-prem or homelab inference servers
  • ✓Teams that must keep code and data off third-party APIs
  • ✓Engineers evaluating whether self-hosting open-weight models beats API costs
  • ✓Hobbyists planning a multi-GPU local LLM rig

Prerequisites

  • •Comfort with Linux administration, GRUB, systemd and Docker
  • •Working knowledge of GPU inference concepts such as VRAM, quantization and tensor parallelism
  • •Hardware-building experience (PCIe, power supplies) and a substantial hardware budget for the top tier

FAQ

What is jamesob's Guide to Running SOTA LLMs Locally?

jamesob's guide to running SOTA LLMs locally is a free, highly detailed GitHub build log for engineers who want to self-host frontier open-weight models. It covers budget tiers from about $2k to $40k+, exact parts, PCIe switch topology, BIOS and kernel settings, power limits and vLLM Docker runners, so you can reproduce a working local inference box.

Is jamesob's Guide to Running SOTA LLMs Locally free?

jamesob's Guide to Running SOTA LLMs Locally is free to access.

What level is jamesob's Guide to Running SOTA LLMs Locally for?

jamesob's Guide to Running SOTA LLMs Locally is aimed at a advanced audience. Recommended background: Comfort with Linux administration, GRUB, systemd and Docker, Working knowledge of GPU inference concepts such as VRAM, quantization and tensor parallelism, Hardware-building experience (PCIe, power supplies) and a substantial hardware budget for the top tier.

How long does jamesob's Guide to Running SOTA LLMs Locally take?

Expect roughly ~1-2 hour read; the full multi-GPU build is a multi-week hardware project. Most learners work through it at their own pace.

What will I learn from jamesob's Guide to Running SOTA LLMs Locally?

You'll learn: Match a hardware budget tier to the open-weight models it can realistically serve; Spec a multi-GPU EPYC workstation with RTX PRO 6000 cards and adequate power; Use an indie PCIe Gen4 switch to get fast GPU peer-to-peer bandwidth; Configure BIOS, GRUB and ACS settings so NCCL multi-GPU P2P stops hanging; Run each model in its own vLLM Docker container with read-only weight mounts; Serve heavily quantized and REAP-pruned frontier MoE models with very long contexts; Power-limit GPUs with nvidia-smi to fit a household electrical circuit safely; Wire a self-hosted coding-agent harness with local search, Git hosting and sandboxing.

Topics

local llmself-hostingvllmmulti-gpuinference hardwareopen-weight models

Sources

This page was written from 3 sources, 2 on domains other than github.com.

  1. 1.github.com — local llmvendor
  2. 2.news.ycombinator.com — item
  3. 3.hn.algolia.com — search