Sail Research
by Sail Research
Inference and sandbox infrastructure built for agents that run for days, not seconds
Sail Research sells inference and compute infrastructure tuned for long-horizon AI agents rather than chat latency. It pairs an OpenAI- and Anthropic-compatible API over open-weight models with Sailboxes, persistent Linux VMs that auto-sleep while an agent waits. It targets engineering teams whose agent token and sandbox bills have outgrown general-purpose GPU clouds.
Sail Research operates an inference and compute platform built specifically for long-horizon AI agents, the workloads that run for hours or days and consume billions of tokens on a single task. The company emerged from stealth on 25 June 2026 with $80 million in combined Seed and Series A funding at a $450 million valuation, backed by Kleiner Perkins, Sequoia Capital, Redpoint, Theory Ventures, Vine Ventures, CRV, A* and Abstract Ventures, plus angels including Alphabet chair John Hennessy, Intel CEO Lip-Bu Tan and Together AI chief scientist Tri Dao. The platform has two halves. The first is an inference stack built on customised open-source engines including vLLM, serving open-weight models such as GLM-5.3, DeepSeek V4 Flash and Pro, Kimi K3 and K2.6, gpt-oss-120b, Gemma 4 and Qwen through OpenAI- and Anthropic-compatible APIs, so migrating requires only a new base URL and key. Its distinguishing feature is completion windows: the caller declares how much latency it will tolerate, Default, Balanced or Flex, and Sail schedules the work across its fleet accordingly, advertising roughly 5-35%, 45-65% and 60-80% savings respectively. The second half is Sailboxes, generally available since 14 July 2026: full Linux VMs with persistent disk, local NVMe, Docker support and no runtime limits, which live-migrate across Sail's fleet based on live resource usage and auto-sleep while an agent waits on inference, so customers pay for work rather than for a reservation. Sail reports 90.72% accuracy on the BrowseComp-Plus research benchmark at up to a tenth of rivals' inference cost, and names Parallel, Detail.dev, Jack and Jill, and Quadrillion Labs as production customers.
The platform or infrastructure engineering lead running agents on open-weight models whose token and sandbox spend has become a line item worth engineering against.
Materially lower cost per agent-hour, from latency-tolerant scheduling on inference plus sandboxes that stop billing while the agent waits.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Enterprise Developers, Platform Engineers, AI/ML Engineers
- Deployment
- API-based, Cloud-only
- Headquarters
- San Francisco, United States
Key Features
- ✓Completion windows
Callers declare latency tolerance per request — Default, Balanced or Flex — and Sail schedules against fleet capacity, cutting token cost by roughly 5-35%, 45-65% or 60-80%.
- ✓Sailboxes
Full Linux VMs with persistent state, independent disks, local NVMe and Docker support that run for hours or days with no runtime limit.
- ✓Auto-sleep and live migration
Sailboxes suspend while an agent waits on inference and migrate across the compute fleet by live resource usage, so idle waiting is not billed as reserved capacity.
- ✓Drop-in compatible APIs
OpenAI Chat Completions, OpenAI Responses and Anthropic Messages endpoints mean migration is a base-URL and API-key change rather than a rewrite.
- ✓Open-weight model catalogue
Serves GLM-5.3 and 5.3-Flash, DeepSeek V4 Flash and Pro, Kimi K3 and K2.6, gpt-oss-120b, Gemma 4 and Qwen with per-model input, cached and output rates.
- ✓Fine-tuning and RL support
Deploys LoRA fine-tunes and supports reinforcement-learning rollout sampling, including a Tinker integration, so custom adapters run on the same stack.
- ✓Zero data retention by default
ZDR is the default posture, with HIPAA and SOC 2 compliance and enterprise BAA, MSA and DPA terms plus regional data processing available.
Capabilities
Use Cases
- •Cost-optimising a deep-research agent
Route a multi-hour browsing and synthesis agent through Flex windows to cut token spend by up to 80% where results are consumed asynchronously.
- •Hosting autonomous coding agents
Give each agent a Sailbox with a real filesystem and Docker so it can build, test and iterate across days without losing working state.
- •Parallel experiment fleets
Fork Sailboxes to run many agent variants concurrently, as Quadrillion Labs does for the Qualia Cloud autonomous research platform.
- •Migrating off a general-purpose inference provider
Swap the base URL on an existing OpenAI- or Anthropic-shaped client and compare cost per token on identical traffic before committing.
- •Serving custom LoRA adapters at agent scale
Deploy domain-tuned adapters and sample RL rollouts on the same inference fleet instead of standing up separate GPU infrastructure.
Ideal For
Best For
- ✓Deep-research and web-browsing agents that run multi-hour task chains and burn billions of tokens
- ✓Batch and offline agent workloads that can trade latency for cost via Flex completion windows
- ✓Long-running coding agents that need a persistent machine with Docker, a real disk and no runtime cap
- ✓Teams already standardised on open-weight models (GLM, DeepSeek, Kimi, Qwen, gpt-oss) looking to cut inference spend
- ✓Reinforcement-learning and LoRA fine-tuning pipelines that need rollout sampling and custom adapter serving
Not Ideal For
- ✗Teams that need frontier proprietary models — Sail serves open-weight models only, so anything built on GPT-5-class or Claude-class models cannot move over
- ✗Latency-critical, human-in-the-loop products, where the Balanced and Flex windows that produce most of the savings are unusable
- ✗Risk-averse enterprises that require a long operating history: Sail left stealth in June 2026 and has no G2, Capterra or TrustRadius presence to check
- ✗Regulated buyers needing on-premises deployment — Sail is an API-based cloud service with no self-hosted option
Integrations
Deployment
Market Analysis
Pros
- ✓Concrete, published per-token rates with a documented mechanism (completion windows) for trading latency against cost
- ✓Sailboxes remove a real cost trap: paying for reserved CPU and memory every second an agent sits waiting on inference
- ✓Drop-in OpenAI and Anthropic API compatibility makes an evaluation cheap — base URL and key, not a rewrite
- ✓Named production customers (Parallel, Detail.dev, Jack and Jill, Quadrillion Labs) and a published BrowseComp-Plus result at 90.72%
Cons
- ✗Open-weight models only — there is no path here for a product built on frontier proprietary models
- ✗Essentially no independent practitioner review surface: no G2, Capterra or TrustRadius listing, and its Hacker News submissions have drawn single-digit points and almost no comments
- ✗The headline savings depend on tolerating latency; workloads that need the Default window keep only the 5-35% band
- ✗The 10x cost advantage rests on engineering optimisation, which The Next Web notes may erode as rivals improve and frontier labs cut prices independently
- ✗Emerged from stealth in June 2026 with Sailboxes GA only since July 2026, so there is no multi-year reliability track record to diligence
Pricing
Self-serve (usage-based)
Pay per token, $5/mo free credits
- ✓Prepaid credits
- ✓DeepSeek V4 Flash at $0.09 input / $0.02 cached / $0.18 output per 1M tokens
- ✓Kimi K3 at $3.00 input / $0.30 cached / $15.00 output per 1M tokens
- ✓Default, Balanced and Flex completion windows
- ✓No strict rate limits
Sailboxes
Usage-based compute
- ✓Full VMs with persistent disk and local NVMe
- ✓Docker support, no runtime limits
- ✓Auto-sleep while agents wait
- ✓Fork for parallel scaling
Enterprise
Contact for pricing
- ✓Custom volume pricing
- ✓HIPAA BAA
- ✓MSA and DPA
- ✓Regional data processing
- ✓SLAs
Metered per million tokens with separate input, cached-input and output rates that vary widely by model — DeepSeek V4 Flash lists at $0.09 input and $0.18 output while Kimi K3 lists at $3.00 and $15.00. The real lever is the completion window: Balanced saves roughly 45-65% and Flex 60-80% against the default ASAP rate, but not every model supports every window. Sailbox compute is billed separately on usage with auto-sleep, and $5 in free credits arrives monthly. Volume discounts, HIPAA BAAs, regional processing and SLAs are Enterprise-only and require a sales conversation.
Security & Compliance
Sources
This page was written from 8 sources, 7 on domains other than sailresearch.com.
- 1.sailresearch.com — sailresearch.comvendor
- 2.docs.sailresearch.com — pricing
- 3.docs.sailresearch.com — docs.sailresearch.com
- 4.prnewswire.com — sail research raises 80 million to build max efficiency infr
- 5.prnewswire.com — sail research launches sailboxes the first cloud environment
- 6.siliconangle.com — sail research raises 80m optimize long horizon ai agents
- 7.thenextweb.com — sail research 80m ai agent inference
- 8.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
DigitalOcean Managed Agents
Managed agent runtime with microVM sandboxes, 16,000+ governed tools and serverless inference, billed on active CPU
Modular
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary
Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters