SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds

by SkyPilot

IntermediateDocumentationFree~1-2 hours to install and land a first GPU job; ongoing as reference

One YAML task spec that runs your training or serving job on whatever GPUs you can actually get.

Start LearningAdded Aug 20, 2026 · Updated Aug 20, 2026

Overview

SkyPilot is an open-source framework for running, managing and scaling AI workloads across any compute you can reach — Kubernetes, Slurm, reserved GPU fleets and more than twenty clouds including AWS, GCP, Azure, OCI, CoreWeave, Nebius, Lambda Cloud, RunPod, Fluidstack, Vast.ai, Crusoe and Prime Intellect — behind a single YAML or Python task definition. It originated at UC Berkeley's Sky Computing Lab, the group also known for vLLM and Ray. The documentation is organised as Getting Started (installation, quickstart, an AI-training tutorial), Running Jobs (distributed jobs, managed jobs, many jobs), Examples (including syncing code and artifacts) and Reference (YAML spec, CLI reference, Python API, async execution, API server). The quickstart moves from a hello_sky.yaml declaring resources such as 'accelerators: A100:8' through 'sky launch', 'sky exec', 'sky status' and 'sky dashboard', then on to managed jobs that provision themselves, auto-recover from spot preemption and tear down when finished. Core capabilities include smart failover across regions and providers, gang scheduling for multi-node training, workload binpacking, autostop for idle clusters, SSH and IDE access into running pods, and — added during 2026 — Endpoints for production inference and Sandboxes for running untrusted agent-generated code on clusters you already own. The project is Apache-2.0 with roughly 10,500 GitHub stars; v0.13.0 shipped in July 2026 with Hugging Face storage, batch inference abstractions and lifecycle hooks.

At a Glance

Topic
Frameworks
Level
Intermediate
Format
Documentation
Cost
Free
Duration
~1-2 hours to install and land a first GPU job; ongoing as reference
Provider
SkyPilot
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • Declare GPU, node and setup requirements once in a portable YAML task spec
  • Launch, exec on, monitor and tear down clusters with the sky CLI
  • Run managed jobs that auto-recover from spot preemption without manual babysitting
  • Use smart failover to find scarce capacity across regions, clouds and Kubernetes
  • Scale to hundreds or thousands of queued runs with the many-jobs workflow
  • Sync code and artifacts between your laptop, object storage and remote clusters
  • Cut idle spend using autostop and binpacking across a shared GPU fleet
  • Drive the same workloads programmatically through the Python API and API server

Highlights

  • One interface over Kubernetes, Slurm and 20+ clouds — no more per-provider launch scripts
  • Comes out of UC Berkeley's Sky Computing Lab, the lab that also produced vLLM and Ray
  • Managed spot jobs with automatic recovery are the headline cost lever, not a footnote feature
  • Apache-2.0 with ~10.5k GitHub stars and v0.13.0 shipped July 2026 — actively developed
  • Recent additions target agent workloads directly: Endpoints for inference, Sandboxes for untrusted code

Who It’s For

Best For

  • ML engineers training or fine-tuning across scarce, multi-provider GPU capacity
  • Platform teams standardising job submission over Kubernetes and cloud VMs
  • Researchers chasing cheap spot GPUs without writing per-cloud automation
  • Teams running batch inference or agent sandboxes on clusters they already own

Prerequisites

  • Working knowledge of Python, YAML and the Linux command line
  • Credentials for at least one cloud account, or access to a Kubernetes or Slurm cluster
  • Familiarity with the GPU training or inference workload you intend to run

FAQ

What is SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds?

SkyPilot is the infrastructure layer for teams who train, fine-tune or serve models on whatever GPUs they can actually get hold of. It presents Kubernetes, Slurm and more than twenty clouds through a single YAML task spec, then handles provisioning, region and provider failover, spot recovery and cleanup. Read these docs if you are tired of maintaining per-cloud launch scripts; afterwards you can run a distributed fine-tune, queue hundreds of managed jobs and serve an inference endpoint without rewriting anything per provider.

Is SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds free?

SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds is free to access.

What level is SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds for?

SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds is aimed at a intermediate audience. Recommended background: Working knowledge of Python, YAML and the Linux command line, Credentials for at least one cloud account, or access to a Kubernetes or Slurm cluster, Familiarity with the GPU training or inference workload you intend to run.

How long does SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds take?

Expect roughly ~1-2 hours to install and land a first GPU job; ongoing as reference. Most learners work through it at their own pace.

What will I learn from SkyPilot Documentation — Run AI Workloads on Kubernetes, Slurm and 20+ Clouds?

You'll learn: Declare GPU, node and setup requirements once in a portable YAML task spec; Launch, exec on, monitor and tear down clusters with the sky CLI; Run managed jobs that auto-recover from spot preemption without manual babysitting; Use smart failover to find scarce capacity across regions, clouds and Kubernetes; Scale to hundreds or thousands of queued runs with the many-jobs workflow; Sync code and artifacts between your laptop, object storage and remote clusters; Cut idle spend using autostop and binpacking across a shared GPU fleet; Drive the same workloads programmatically through the Python API and API server.

Topics

GPU orchestrationmulti-cloudspot instancesKubernetesSkyPilot

Sources

This page was written from 3 sources, 2 on domains other than docs.skypilot.ai.

  1. 1.docs.skypilot.ailatestvendor
  2. 2.github.comskypilot
  3. 3.hn.algolia.comhn.algolia.com