SkyPilot
by SkyPilot Inc.
Run and scale AI workloads across Kubernetes, Slurm and 25+ clouds from one interface
SkyPilot is an open-source AI compute platform that unifies Kubernetes clusters, Slurm clusters, on-premises machines and more than 25 cloud providers into a single pool, so AI and platform teams can launch training, fine-tuning, batch inference and serving workloads anywhere GPUs are actually available instead of hand-managing each environment separately.
SkyPilot is an AI compute platform that presents Kubernetes clusters, Slurm clusters, SSH-accessible on-premises machines and more than 25 cloud providers as a single pool of infrastructure, so an AI team can launch a training run, a fine-tuning job, a batch inference sweep or an agent sandbox with one YAML file or one Python call and let the system decide where it actually lands. It began as an open-source project in the UC Berkeley Sky Computing Lab, the same lab behind Spark and Ray, and the Apache-2.0 repository carries roughly 10,400 GitHub stars, with v0.12.0 in March 2026 adding Slurm support, job groups and improved data mounting. The core abstractions are clusters (virtual collections of VMs or pods you can SSH into and attach an IDE to), jobs (sky exec on an existing cluster, or sky jobs launch for managed jobs that auto-recover and fail over to another region or provider when capacity disappears) and services for multi-replica model serving spread across locations and pricing tiers. SkyPilot Inc. came out of stealth on 21 July 2026 with a $20 million seed round led by Lux Capital, alongside a commercial SkyPilot Platform that layers gang scheduling, priority queues, GPU sharing and quotas, GPU health checks with auto-remediation, sub-second inference sandboxes, SSO, RBAC, cost reporting and a unified dashboard on top of the open-source core. The vendor lists Nubank, Meta, NVIDIA, Amazon Robotics, Shopify, Mistral, Abridge, HeyGen and Hippocratic AI among more than 30 enterprises using it.
The platform or ML infrastructure team at an organisation whose GPU capacity is split across a Kubernetes cluster, a Slurm cluster and two or more clouds or neoclouds, and who is currently maintaining a separate launch path for each one.
One YAML or Python interface that places every AI job on whichever of your existing clusters and cloud accounts has capacity, with automatic failover when a region runs out of GPUs.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Freemium, Contact for pricing
- Target Market
- CTOs, VPs of Engineering, Platform Engineers, ML Infrastructure Teams, Data Scientists, Enterprise Developers
- Deployment
- Open-source, Self-hosted, Hybrid, Multi-cloud
- Headquarters
- San Francisco, United States
- Customers
- 30+ enterprises per the vendor, including Nubank, Meta, NVIDIA, Amazon Robotics, Shopify, Mistral, Abridge, HeyGen and Hippocratic AI
Key Features
- ✓Unified multi-infrastructure interface
Presents Kubernetes, Slurm, SSH node pools and 25+ clouds as one compute pool addressed through a single YAML or Python API.
- ✓Managed jobs with auto-recovery and failover
sky jobs launch retries and relocates a job to another region or provider when nodes are preempted or capacity disappears mid-run.
- ✓Advanced scheduling
Gang scheduling, multi-node jobs, priority queueing, GPU sharing, quotas and workload binpacking to raise utilisation on fleets you already pay for.
- ✓GPU Manager health checks
Commercial platform monitors accelerator health and auto-remediates bad nodes, which is the failure mode that silently wastes large training runs.
- ✓Services for model serving
Runs multi-replica inference endpoints spread across regions and pricing tiers, with autoscaling and failover between them.
- ✓Sandboxes for agent code
Launches isolated inference and agent execution sandboxes on your own Kubernetes in under a second, per the vendor.
- ✓Team governance controls
SSO, RBAC, policy enforcement, per-team cost reporting and a shared dashboard via a team-wide API server deployment.
Capabilities
Use Cases
- •Chasing scarce GPU capacity across providers
Submit one job definition and let SkyPilot find available H100 or B200 capacity across clouds, regions and neoclouds rather than manually polling each console.
- •Fault-tolerant pre-training and RL runs
Run multi-day distributed training on spot or preemptible capacity, with managed jobs restarting and relocating the run automatically after node loss.
- •Consolidating Kubernetes and Slurm
Give research users one submission path while the platform team keeps central scheduling, quota and cost visibility across both orchestrators.
- •Raising utilisation on an existing GPU fleet
Apply binpacking, priority queues and idle-resource cleanup to squeeze more throughput from committed capacity, which the CEO frames as a 10%-plus utilisation gain.
- •Portable batch inference and evaluation sweeps
Fan a large evaluation or embedding job out across whichever clusters and cloud accounts are cheapest and idle at the time.
Ideal For
Best For
- ✓Multi-cloud and neocloud GPU fleets where capacity has to be chased across AWS, GCP, Azure, CoreWeave, Nebius, Lambda and RunPod
- ✓Long-running pre-training and post-training or reinforcement learning jobs that need automatic recovery when a spot instance or node dies
- ✓Platform teams consolidating Kubernetes and Slurm under a single scheduling and quota layer
- ✓Batch inference sweeps that need to burst beyond the capacity of one cluster or one cloud region
- ✓Research and AI teams that want SSH and IDE access to remote GPUs without writing per-cloud provisioning scripts
Not Ideal For
- ✗Teams entirely standardised on one cloud's managed AI service, such as Amazon SageMaker or Google Vertex AI, where a portability abstraction adds a layer without removing a problem
- ✗Organisations without Kubernetes or cloud operations skills in-house: SkyPilot orchestrates infrastructure you already own and does not hide the underlying cluster, quota or networking work
- ✗Buyers looking for GPU supply rather than GPU orchestration. SkyPilot brokers no capacity, so you still negotiate every cloud contract, reservation and quota increase yourself
- ✗Classic BI, ETL or non-accelerated analytics workloads, which are better served by conventional data orchestration tools
Integrations
Deployment
Market & Ratings
30+ enterprises per the vendor, including Nubank, Meta, NVIDIA, Amazon Robotics, Shopify, Mistral, Abridge, HeyGen and Hippocratic AI
Market Analysis
Pros
- ✓Vendor neutrality is real and structural, not marketing: the same job definition runs on Kubernetes, Slurm and 25+ clouds, which is genuine leverage in GPU contract negotiations
- ✓Mature open-source core with four years of history, 10,400 GitHub stars and active releases, so the technology is not new even though the company is
- ✓Managed jobs with automatic failover directly address the most expensive real-world failure mode, losing a multi-day training run to a preempted node
- ✓Adoption can start free and self-hosted with no procurement cycle, then move to the commercial platform when SSO, RBAC and quotas are needed
Cons
- ✗The commercial platform only came out of stealth on 21 July 2026, so enterprise references, support SLAs and roadmap stability are unproven; lead investor Brandon Reeves publicly described the product as 'probably like 1% of the way done'
- ✗The GitHub tracker carries roughly 118 open issues clustering on provider-specific breakage (Azure, Nebius, Vast), SSH and passwordless-sudo friction on on-prem and Slurm, and fractional-GPU requests breaking the managed-jobs dashboard
- ✗Security defaults have drawn issues of their own, including Azure networking configuration that leaves systems exposed and API key exposure risks, so a hardening pass is required rather than optional
- ✗No published pricing for SkyPilot Platform and no self-serve path, so evaluating cost means entering a sales cycle
- ✗No presence on G2, Capterra or TrustRadius and only light Hacker News discussion since the original 2022 Berkeley launch thread, so there is little independent buyer evidence to weigh against the vendor's own customer list
Pricing
Open Source
$0
- ✓Apache-2.0 licence
- ✓CLI and Python SDK
- ✓Kubernetes, Slurm, SSH node pools and 25+ clouds
- ✓Managed jobs with auto-recovery
- ✓Services for model serving
- ✓Self-hosted API server for team sharing
SkyPilot Platform
Contact for pricing
- ✓Intelligent and gang scheduling, priority queues, quotas
- ✓GPU health checks with auto-remediation
- ✓Sandboxes and endpoints
- ✓SSO, RBAC and policy enforcement
- ✓Cost reporting and unified dashboard
- ✓BYOC/BYOK, private VPC and airgapped deployment
The open-source core is Apache-2.0 and free to self-host with no seat or node fee. SkyPilot Platform publishes no list price at all: the site routes every commercial enquiry to a demo booking, so expect a negotiated enterprise agreement. Either way SkyPilot bills nothing for compute itself. You keep paying your clouds, neoclouds and Kubernetes providers directly, and the platform's pitch is that better utilisation of that existing spend more than covers its own cost.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 7 on domains other than skypilot.ai.
- 1.siliconangle.com — skypilot nabs 20m ease ai infrastructure management
- 2.fortune.com — skypilot from databricks cofounder raises 20m to be the swit
- 3.github.com — skypilot
- 4.github.com — issues
- 5.hn.algolia.com — search
- 6.hpcwire.com — skypilot launches with 20m to accelerate custom intelligence
- 7.skypilot.ai — skypilot.aivendor
- 8.docs.skypilot.ai — overview
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
MyDecisive
Open-source, OpenTelemetry-native telemetry control plane that acts on incidents instead of just charting them
General Compute
ASIC-based inference cloud built for agents, where latency compounds
Cisco Cloud Control
One console where humans and AI agents run and defend the whole Cisco estate
Vector Core Compute
Disaggregated enterprise inference cloud running CPUs, GPUs and RDUs in one pipeline