DeepInfra
by DeepInfra
Purpose-built inference cloud serving 200+ open-source AI models on OpenAI-compatible APIs
DeepInfra is a cloud platform built specifically for high-throughput AI inference, giving engineering and platform teams production access to 200+ open-source models through OpenAI-compatible APIs. It owns and operates its own GPU fleet across nine data centers and is SOC 2 and ISO 27001 certified with a zero data retention policy.
DeepInfra, founded in 2022 and headquartered in Palo Alto, California, operates an inference cloud engineered from the ground up for model serving rather than general-purpose compute. The platform exposes more than 200 open-source models — spanning text generation, embeddings, automatic speech recognition and text-to-image, including Llama, Mistral and Anthropic Claude offerings — through OpenAI-compatible APIs, so teams can migrate existing integrations with minimal code change. DeepInfra processes close to five trillion tokens weekly and says token volume has grown 25x since its Series A, driven by the continuous high-volume demand of agentic AI systems. Unlike resellers, DeepInfra owns and operates its GPU infrastructure: on 8 July 2026 it opened its ninth data center and first international site in Toronto, a 1.7 MW facility hosting over 1,000 NVIDIA Blackwell B300 GPUs, expanding low-latency capacity beyond its eight U.S. sites. CEO and co-founder Nikola Borisov framed the expansion around enterprises "moving from experimentation to production at unprecedented speed." The company closed a $107 million Series B on 4 May 2026, co-led by 500 Global and Georges Harik, with participation from NVIDIA, Samsung Next, Supermicro, Felicis, A.Capital Ventures, Crescent Cove, Peak6 and Upper90. Pricing is usage-based: language models bill per million input and output tokens (for example, Meta-Llama-3.1-8B-Instruct at $0.02/$0.04), most other models bill on inference execution time, and dedicated deployments are billed hourly per GPU (A100 80GB at $0.89, H100 80GB at $2.20, B200 180GB at $3.69). Three scheduling tiers — Flex at 0.8x, Standard at 1x and Priority at 1.5x base rates — let teams trade latency against cost. DeepInfra is SOC 2 and ISO 27001 certified with zero data retention, and collaborates with NVIDIA on Nemotron models and Blackwell deployment.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based
- Target Market
- CTOs, Platform Engineering Leaders, Enterprise Developers, Data Scientists, MLOps Engineers
- Founded
- 2022
- Headquarters
- Palo Alto, United States
Key Features
- ✓200+ open-source models on one API
Serves text generation, embeddings, speech recognition and text-to-image models including Llama, Mistral and Claude through a single endpoint.
- ✓OpenAI-compatible APIs
Existing OpenAI SDK integrations can point at DeepInfra with minimal code changes, lowering migration cost.
- ✓Owned GPU fleet across nine data centers
DeepInfra owns and operates its own infrastructure, including a 1.7 MW Toronto site with 1,000+ NVIDIA Blackwell B300 GPUs opened July 2026.
- ✓Dedicated GPU deployments
Custom LLM deployments bill hourly per GPU — A100 80GB at $0.89, H100 80GB at $2.20 and B200 180GB at $3.69 per GPU-hour.
- ✓Flex, Standard and Priority scheduling tiers
Three service tiers priced at 0.8x, 1x and 1.5x base rates let teams trade latency guarantees against cost.
- ✓Enterprise security posture
SOC 2 and ISO 27001 certified with a zero data retention policy on inference traffic.
Capabilities
Use Cases
- •Production LLM serving at scale
Runs high-throughput text generation workloads that collectively process close to five trillion tokens per week.
- •Agentic AI backends
Supplies the continuous, high-volume inference that autonomous agent systems generate in production.
- •Multimodal AI pipelines
Provides embeddings, automatic speech recognition and text-to-image inference alongside text generation on one platform.
Ideal For
Best For
- ✓Serving open-source LLMs in production without managing GPU infrastructure
- ✓High-volume agentic workloads that need predictable per-token inference costs
- ✓Migrating off proprietary model APIs using an OpenAI-compatible drop-in endpoint
Integrations
Market Analysis
Pros
- ✓Very low per-token pricing versus proprietary model APIs
- ✓Owned infrastructure gives capacity and pricing control others lack
- ✓Drop-in OpenAI compatibility keeps migration cost low
- ✓Zero data retention plus SOC 2/ISO 27001 suits regulated buyers
Cons
- ✗No free tier or trial listed in pricing documentation
- ✗Focused on open-source model serving — not an end-to-end MLOps or agent platform
- ✗Cloud-only; no on-premise or hybrid deployment option
- ✗Geographic footprint is still heavily U.S.-weighted despite the Toronto expansion
Pricing
Serverless inference (per token)
From $0.02/1M input tokens
- ✓Per-million input and output token billing
- ✓200+ open-source models
- ✓OpenAI-compatible API
- ✓Example: Meta-Llama-3.1-8B-Instruct at $0.02/$0.04 per 1M tokens
Dedicated GPU deployment
From $0.89/GPU-hour
- ✓A100 80GB at $0.89/GPU-hour
- ✓H100 80GB at $2.20/GPU-hour
- ✓B200 180GB at $3.69/GPU-hour
- ✓Custom LLM deployments
Scheduling tiers
0.8x-1.5x base rate
- ✓Flex at 0.8x base
- ✓Standard at 1x base
- ✓Priority at 1.5x base
Language models bill per million input/output tokens; most other model types bill on inference execution time. Monthly invoicing thresholds span $20-$10,000 across five usage tiers. No free tier or trial is listed in the pricing documentation.
Sources
This page was written from 3 sources, 1 on domains other than deepinfra.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
WEKA NeuralMesh
Microservices storage and memory fabric built for AI training and inference at exabyte scale
Tsuga
Bring-your-own-cloud observability that keeps telemetry, and its cost, inside your own AWS account
OpenObserve
Open-source, Rust-based observability on object storage — logs, metrics, traces and LLM telemetry in one binary
Coralogix
AI-native observability that queries logs, metrics and traces in place — index-free, in your own S3 bucket