Replicate
by Replicate (acquired by Cloudflare)
Run any of 50,000+ open-source AI models with one API call and per-second billing
Replicate is a hosted API for running open-source machine learning models without provisioning GPUs. Developers call a community model in a line of code, fine-tune it on their own data, or package a proprietary model with the open-source Cog container format and deploy it on Replicate's infrastructure, paying per second of compute rather than for idle capacity.
Replicate is a cloud API that runs open-source machine learning models so developers never have to provision a GPU, build a serving stack or write a Dockerfile for CUDA. A prediction is a single HTTP call or a few lines of Python or Node, and the platform handles scheduling, scaling to zero and returning the output, billed per second of the hardware actually used. Underneath sits Cog, Replicate's open-source container format for packaging models, which is what lets the same interface serve a community image model, a customer's private fine-tune and a large language model without bespoke integration work in each case. Cloudflare's November 2025 acquisition announcement put the catalogue at more than 50,000 production-ready models, spanning image and video generation, speech and music synthesis, image restoration, captioning and text models. Around that core the product adds fine-tuning — training a custom FLUX or SDXL variant on a user's own images — plus Deployments for pinning dedicated capacity to a model, private models, webhooks for asynchronous jobs, organisation accounts, LoRA support and documented rate limits and data-retention policies. Pricing is published in full and unusually granular for the category: CPU from $0.000025 per second, an Nvidia T4 at $0.000225 per second, an L40S at $0.000975, an A100 80GB at $0.0014 and an H100 at $0.001525, with multi-GPU configurations scaling proportionally and some models billed per token or per generated image instead. Companies including Labelbox, Character.ai, Photo.ai, HeadshotPro and Magnific are named as building on the platform. Cloudflare announced the acquisition on 17 November 2025, with the catalogue to be folded into Workers AI.
A product engineering team that wants a specific open-source model in production this week and has no interest in owning GPU capacity planning, container builds or an inference server.
A working, autoscaling inference endpoint for any of 50,000+ models in a single API call, billed per second, with no idle GPU cost between requests.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Enterprise Developers, Data Scientists, ML Engineers, Startup Founders
- Deployment
- Cloud-only, API-based
Key Features
- ✓50,000+ model catalogue
A community library spanning image, video, audio, language and restoration models, all callable through one consistent prediction API.
- ✓Cog packaging format
An open-source container standard for ML models that makes a custom model deployable with the same interface as any community model.
- ✓Per-second hardware billing
Published rates from $0.000025/sec on CPU to $0.001525/sec on an H100, so idle time between requests costs nothing.
- ✓Fine-tuning
Train a custom variant of models such as FLUX or SDXL on your own dataset and serve it through the same endpoint.
- ✓Deployments
Pin dedicated capacity to a specific model version to control scaling behaviour and reduce cold-start latency for production traffic.
- ✓Webhooks and async predictions
Long-running generations report back by webhook rather than holding an HTTP connection, which is essential for video and large image jobs.
- ✓Private models and organisations
Keep proprietary weights private and manage access, API tokens and billing across a team from one account.
Capabilities
Use Cases
- •Adding image generation to a SaaS product
Call a hosted diffusion model from application code and ship a generation feature without operating any GPU infrastructure.
- •Custom fine-tuned avatars and headshots
Train a per-customer model on uploaded photos, then serve each user's private fine-tune through the same prediction API.
- •Model evaluation and bake-offs
Run the same prompt across a dozen competing open-source models to pick one, paying only for the seconds consumed.
- •Speech, music and audio pipelines
Chain transcription, synthesis and enhancement models into a pipeline without managing separate serving stacks for each.
- •Deploying a proprietary in-house model
Package a private model with Cog and get autoscaling inference plus webhooks without building a serving layer.
Ideal For
Best For
- ✓Shipping image, video, speech or music generation features without hiring an ML infrastructure engineer
- ✓Prototyping across many open-source models quickly, where the switching cost between models is one string
- ✓Spiky or unpredictable inference traffic where reserved GPU capacity would sit idle most of the time
- ✓Fine-tuning image models such as FLUX or SDXL on customer data and serving the result through the same API
- ✓Packaging and deploying a proprietary model with Cog without building a serving stack in-house
Not Ideal For
- ✗High, steady-volume inference — practitioners on Hacker News put serverless GPU markups at more than 10x on-demand pricing, so at sustained load reserved instances or self-hosting win on cost
- ✗Latency-critical interactive workloads on infrequently-used models, because cold boots are a known weakness the founder has publicly acknowledged
- ✗Teams with strict data-residency or on-premise requirements, since this is a cloud-only API with no self-hosted control plane
- ✗Buyers who need certainty about long-term platform direction while the Cloudflare integration into Workers AI is still in progress
Integrations
Deployment
Market Analysis
Pros
- ✓Time to a working inference endpoint is minutes, not the weeks a self-hosted serving stack takes
- ✓Per-second billing with scale-to-zero means no cost between requests, which suits bursty product traffic
- ✓Pricing is published in full down to the per-second hardware rate, which is rare in this category
- ✓Cog is open source, so a model packaged for Replicate is not trapped there
Cons
- ✗Cold starts are a real and acknowledged weakness — co-founder Ben Firshman replied directly on Hacker News: 'Yeah, our cold boots suck', while outlining mitigations such as Deployments and image-distribution work
- ✗Cost at scale is the standard objection: one practitioner argued serverless GPU markups 'compared to on-demand pricing is usually >10x', recommending spot instances or self-hosted clusters for cost-sensitive volume
- ✗No free tier at all, so evaluation costs money from the first prediction
- ✗Cloud-only with no self-hosted or in-VPC option, which rules it out for workloads that cannot send data to a third party
- ✗The Cloudflare acquisition has created visible uncertainty about independent direction — an August 2026 Hacker News post asked outright whether there is 'still a functioning team behind replicate.com'
Pricing
Pay-as-you-go (CPU)
From $0.09/hr
- ✓CPU Small at $0.000025/sec ($0.09/hr)
- ✓Standard CPU at $0.000100/sec ($0.36/hr)
- ✓Billed per second of run time
- ✓Prepaid credits supported
Pay-as-you-go (GPU)
From $0.81/hr
- ✓Nvidia T4 at $0.000225/sec ($0.81/hr)
- ✓Nvidia L40S at $0.000975/sec ($3.51/hr)
- ✓Nvidia A100 80GB at $0.001400/sec ($5.04/hr)
- ✓Nvidia H100 at $0.001525/sec ($5.49/hr)
- ✓2x, 4x and 8x multi-GPU configurations scale proportionally
Enterprise / volume
Contact for pricing
- ✓Volume discounts
- ✓Dedicated account manager
- ✓Priority support
- ✓Higher GPU limits
- ✓Performance SLAs
Billing is per second of the hardware a prediction occupies, with every rate published openly — $0.000225/sec for a T4 up to $0.001525/sec for an H100 — and some models metered per input token, output token or generated image instead. Fast-booting fine-tunes are charged only for active processing, not idle time. There is no free tier; all usage is paid. Volume discounts, SLAs and a dedicated account manager sit behind an enterprise conversation. The economics favour spiky traffic: practitioners report the serverless premium runs well above on-demand GPU pricing at sustained load.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 3 on domains other than replicate.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Xiaomi MiMo-V2.6
MIT-licensed, 1M-context multimodal MoE models that took the top open-weights score on Artificial Analysis at well under $1 per million tokens
TypeSafe Jev
A 'System One' decision model that returns typed choices, scores and calibrated probabilities instead of text, for agent routing and classification
PrismML Bonsai
Open-weight 1-bit and ternary LLMs that run 27B-class AI on laptops and phones
Sakana AI Fugu Max
Multi-agent orchestration behind one API — frontier-grade results at $2 per million input tokens