Parasail
by Parasail
Pay-per-token inference cloud for open models, with no long-term GPU contracts.
Parasail is a pay-per-token inference cloud for open-weight models. It is built for AI-native startups and engineering teams that need production LLM endpoints without long-term GPU contracts. One OpenAI-compatible API covers serverless, dedicated GPU, workload-tuned elastic and batch inference, with work spread across Parasail's own GPUs and partner GPU capacity.
Parasail is a San Francisco inference cloud that sells open-model LLM inference on a pay-per-token basis, with no long-term GPU contracts. The company came out of stealth in April 2025. Its founder and CEO, Mike Henry, previously built Groq's cloud offering. Instead of running a single GPU fleet, Parasail runs what it calls an AI Supercloud. This orchestration layer combines Parasail's own GPUs with rented capacity from outside providers across roughly 40 data centers in 15 countries, according to TechCrunch and SiliconANGLE. It also automates kernel configuration, quantization and endpoint tuning, so developers can deploy with a few lines of code. Buyers get four products behind one OpenAI-compatible API at api.parasail.io/v1. Serverless endpoints are billed per token; for example, DeepSeek V4 Flash costs $0.14 per million input tokens and $0.28 per million output tokens. Elastic Endpoints, still in early access, are private endpoints tuned for each workload. Dedicated deployments run on reserved NVIDIA B300, B200, H200, H100, RTX PRO 6000 or RTX 5090 GPUs with negotiated SLAs. An OpenAI-format batch API is priced at 50% off serverless, according to the docs. Parasail is also a provider on OpenRouter, where it lists 39 models from DeepSeek, Qwen, MoonshotAI, Z.ai, Meta, Mistral and others. In April 2026 it raised a $32 million Series A co-led by Touring Capital and Kindred Ventures, bringing total funding to $42 million. At that time it said it served 500 billion tokens a day; its homepage now claims trillions of tokens served daily. Customers named in the funding announcement include Elicit, mem0, Gravity, Kotoba and Venice. Parasail targets AI-native startups and competes with better-funded inference specialists such as Fireworks AI and Baseten.
Engineering leads at AI-native startups who run open-weight LLMs in production and want per-token pricing, with dedicated GPUs available, without signing long-term GPU contracts.
Production open-model inference behind one OpenAI-compatible API, with batch jobs priced at half the serverless rate.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Enterprise Developers, ML Engineers, AI-native startups
- Deployment
- Cloud-only, API-based
- Headquarters
- San Francisco, USA
Key Features
- ✓Serverless endpoints
Pay-as-you-go access to popular open models with no minimums or setup, billed per million input and output tokens.
- ✓Dedicated deployments
Reserved NVIDIA B300, B200, H200, H100 or RTX-class GPUs sized to your roadmap, with negotiated SLAs and private control of the model.
- ✓Batch processing API
OpenAI-compatible batch API for offline jobs of millions of requests, priced at 50% off serverless according to the docs.
- ✓Elastic Endpoints (early access)
Private endpoints tuned for each workload and priced per million tokens, with capacity that scales with actual traffic instead of reserved GPUs.
- ✓OpenAI-compatible API
Point an existing OpenAI SDK at api.parasail.io/v1 with a bearer API key, so moving from another provider needs minimal code changes.
- ✓AI Supercloud orchestration
Routes workloads across Parasail's own GPUs and outside providers in about 40 data centers, and automates kernel and quantization tuning.
- ✓Commit-based drawdown billing
Prepaid commitments can be spent on any model, instead of locking spend to one GPU reservation or one model.
Capabilities
Use Cases
- •Agent backends on open models
Startups building agent products on models such as Kimi K3, GLM-5.3 or DeepSeek V4 Flash get tool-calling endpoints without buying GPU capacity.
- •Bulk document summarization
Publishers and research tools summarize large archives through the batch API. SiliconANGLE cites a scientific publisher summarizing its archive of academic papers this way.
- •Serving a fine-tuned custom model
Teams deploy their own fine-tuned open model to a dedicated endpoint, optionally FP8-quantized to halve its memory needs.
- •Retrieval-augmented generation
Engineering teams build RAG pipelines on hosted embedding and chat models behind the same OpenAI-compatible API, with no separate vendor integration.
- •Reducing inference spend
Cost-sensitive teams move open-model traffic to per-token and discounted batch pricing to cut inference bills without signing long-term capacity contracts.
Ideal For
Best For
- ✓AI-native startups serving open-weight LLMs in production without committing to GPU contracts
- ✓Large offline batch inference jobs where cost per token matters more than latency
- ✓Teams already on the OpenAI SDK that want to switch to open models with minimal code changes
- ✓Hosting custom or fine-tuned open models on dedicated GPUs with negotiated SLAs
Not Ideal For
- ✗Regulated enterprises that need SOC 2, ISO 27001 or HIPAA attestations up front: the Parasail trust center page we reviewed describes internal security policies but states no certification for Parasail itself
- ✗Teams that want model training alongside serving: TechCrunch describes Parasail as inference-only, so training needs a separate provider
- ✗Buyers who need published list prices for dedicated GPUs: hourly rates are available only on request
Deployment
Market Analysis
Pros
- ✓Publishes per-token serverless and batch prices, and batch is priced at 50% off serverless
- ✓The OpenAI-compatible API and its listing as an OpenRouter provider (39 models) make switching cheap
- ✓Serverless, elastic, dedicated and batch inference all sit under one account, with no long-term GPU contracts
- ✓Funded with $42M total, and the CEO has previously built an inference cloud (Groq's)
- ✓A Hacker News commenter (May 2025) independently named Parasail as a source for open-model batch inference about 50% cheaper than live inference
Cons
- ✗Young vendor that came out of stealth only in April 2025. TechCrunch notes its customers are concentrated among seed to Series B AI startups and that rivals Fireworks AI and Baseten are better funded
- ✗An independent GitHub audit of OpenRouter providers lists, as an unchecked lead, that Parasail's pricing separates FP8 and FP16 tiers without saying which serverless models run at which precision. Parasail's own FP8 docs cover dedicated deployments only
- ✗Little public compliance evidence: the trust center page we reviewed lists internal policies but no SOC 2, ISO 27001 or HIPAA attestation for Parasail itself
- ✗Dedicated GPU hourly rates and Elastic Endpoint prices are not published, and Elastic Endpoints are still in early access
- ✗Few independent user reviews: Hacker News has only passing comments, Reddit searches found nothing relevant, and G2 could not be accessed
Pricing
Serverless Endpoints
From $0.09 per 1M input tokens
- ✓No minimums
- ✓Per-token pricing varies by model
- ✓Community support
Batch Processing
From $0.02 per 1M input tokens
- ✓Millions of requests per job
- ✓Tiered by model size and precision (FP4/FP8/FP16); the $0.02 rate is FP4 for 0-4.1B models
- ✓OpenAI-compatible batch API
Elastic Endpoints (early access)
Contact for pricing
- ✓Private optimized endpoints priced per million tokens
- ✓Tuning for each workload
- ✓Dedicated support
Dedicated Deployments
Contact for pricing
- ✓Reserved NVIDIA B300/B200/H200/H100/RTX PRO 6000/RTX 5090 GPUs
- ✓Negotiated SLAs
- ✓Volume discounts
Serverless is pay-as-you-go per million tokens with no minimums. Examples: DeepSeek V4 Flash costs $0.14 input and $0.28 output; Mistral Small 3.2 24B costs $0.09 and $0.30; Qwen3-Next 80B costs $0.10 and $1.10. Batch pricing is tiered by model size and precision, from $0.02 input and $0.04 output for 0-4.1B models in FP4 up to $0.97 input and $2.51 output for 500B+ models in FP16; the docs describe batch as 50% off serverless. Elastic Endpoints (early access) are priced per million tokens through sales. Dedicated GPU hourly rates are available only on request, with volume discounts. Prepaid commitments can be spent on any model. No free tier or free trial was confirmed.
Security & Compliance
Connect
Sources
This page was written from 12 sources, 10 on domains other than parasail.io.
- 1.parasail.io — parasail.iovendor
- 2.parasail.io — pricingvendor
- 3.docs.parasail.io — parasail docs
- 4.docs.parasail.io — fp8 quantization
- 5.trust.parasail.io — trust.parasail.io
- 6.techcrunch.com — parasail raises 32m to feed tokenmaxxing ai developers
- 7.siliconangle.com — parasail raises 32m pay per token inference cloud
- 8.prnewswire.com — parasail raises 32m series a to build the supercloud that pu
- 9.kindredventures.com — parasail the ai supercloud for the agent era
- 10.openrouter.ai — parasail
- 11.github.com — provider transparency.md
- 12.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
DigitalOcean Managed Agents
Managed agent runtime with microVM sandboxes, 16,000+ governed tools and serverless inference, billed on active CPU
Modular
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary
Lambda
GPU cloud and AI factories for training and inference — on-demand NVIDIA instances to single-tenant superclusters