Prime Intellect
by Prime Intellect
Rent GPUs, post-train agents with RL, and serve your own models on one open stack.
Prime Intellect is a full-stack platform for training, evaluating and serving your own AI models: a multi-provider GPU marketplace, hosted reinforcement-learning post-training through its Lab product, a community Environments Hub, and inference. It is for AI engineering teams that want custom agentic models without depending on a frontier lab's closed API.
Prime Intellect is a San Francisco company, founded in January 2024 by Vincent Weißer and Johannes Hagemann, that sells the infrastructure for training and serving your own AI models rather than renting a frontier lab's API. It began in July 2024 as a compute exchange that aggregates GPU capacity from many clouds; today its homepage lists on-demand H100, H200, B200 and B300 instances, spot capacity and reserved InfiniBand clusters with Slurm or Kubernetes orchestration. On top of that sits Lab, launched from private beta on February 10, 2026 after more than 3,000 RL runs, which bundles hosted reinforcement-learning training, evaluations, sandboxes for code execution and inference (dedicated, LoRA pay-per-token and serverless OpenAI-compatible) around the Environments Hub, a community registry of RL environments opened in August 2025. The software underneath is open source: prime-rl, an Apache-2.0 asynchronous RL trainer built on FSDP2 and vLLM, and verifiers, an MIT library for environments and evals. Its model line runs from INTELLECT-1 (10B, November 2024) and INTELLECT-2 (32B, May 2025) to INTELLECT-3, a 106B mixture-of-experts released in November 2025 and post-trained from Z.ai's GLM-4.5-Air on 512 H200 GPUs across 64 nodes, a centralized cluster despite the company's decentralized-training origins. In July 2026 it raised a $130 million Series A led by Radical Ventures at a reported $1 billion valuation, disclosing an annualized revenue run rate above $100 million and customers including Ramp and Zapier. It competes with GPU clouds such as Lambda and RunPod and with managed fine-tuning platforms such as Together AI.
Heads of ML or AI platform teams at AI-native companies who want to post-train and serve their own agentic models on open weights instead of paying per call for a frontier API.
RL-tuned custom models trained, evaluated and deployed on one platform, with GPU capacity bought per minute across many providers.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, ML Engineers, AI Researchers, Enterprise Developers
- Deployment
- Multi-cloud, Open-source, API-based
- Founded
- 2024
- Headquarters
- San Francisco, United States
- Customers
- 6,000+ (company-reported, July 2026)
Key Features
- ✓Multi-provider GPU marketplace
Rents on-demand, spot and reserved NVIDIA GPUs sourced from many providers through one account, so teams compare prices without negotiating separate cloud contracts.
- ✓Lab hosted RL training
Runs reinforcement-learning post-training, evaluations and inference as a managed service, removing the need to operate your own GPU cluster for agent training.
- ✓Environments Hub
A community registry of 2,500+ RL environments built on the verifiers library, reusable both as training tasks and as benchmark evaluations.
- ✓prime-rl open-source trainer
Apache-2.0 asynchronous RL framework using FSDP2 and vLLM that scales from a single GPU to 1,000+ GPUs, so workloads are portable off the platform.
- ✓Code-execution sandboxes
Secure remote sandboxes execute model-generated code during training and evaluation, which coding and tool-use environments need to score model outputs.
- ✓Inference deployment
Serves trained models through dedicated deployments, pay-per-token LoRA adapters or serverless OpenAI-compatible endpoints, closing the loop from training to production.
- ✓Multi-node clusters
InfiniBand-connected clusters with Slurm or Kubernetes orchestration and Grafana monitoring support large distributed pre-training and RL jobs.
Capabilities
Use Cases
- •Train a task-specific agent
Post-train an open-weight model with RL on your own environment to outperform general frontier models on a narrow task; Ramp reports its spreadsheet-analysis agent did so faster and cheaper.
- •Burst GPU capacity for experiments
Rent a single H100 or a discounted spot instance for short fine-tuning or evaluation jobs, paying per minute of runtime instead of committing to a reserved contract.
- •Model evaluation before selection
Run hosted evaluations of open models against Hub environments to compare candidates before choosing which base model to invest post-training effort in.
- •Large-scale distributed training
Reserve InfiniBand multi-node clusters with Slurm or Kubernetes for multi-week pre-training or RL runs that outgrow a single eight-GPU node.
- •Reduce frontier API dependency
Replace a closed-model API call with a smaller custom model served on dedicated or serverless inference, gaining control over weights, per-request cost and data handling.
Ideal For
Best For
- ✓AI engineering teams post-training open-weight models with reinforcement learning for agentic tasks
- ✓Startups and research labs that need flexible, per-minute GPU rental across multiple providers
- ✓Teams building RL environments and evals they want to reuse across training and benchmarking
- ✓Organizations that want an open-source training stack (prime-rl, verifiers) they can also run on their own hardware
Not Ideal For
- ✗Regulated enterprises that need documented SOC 2, HIPAA, SSO or data-residency commitments: none are published in the docs FAQ or on the homepage
- ✗Workloads that require contractual uptime guarantees, since the FAQ states there are no formal SLAs because capacity comes from third-party providers
- ✗Business teams looking for an off-the-shelf AI application; this is infrastructure that assumes in-house ML engineering skills
Integrations
Deployment
Market & Ratings
6,000+ (company-reported, July 2026)
Market Analysis
Pros
- ✓Open-source core (prime-rl under Apache-2.0, verifiers under MIT) reduces lock-in
- ✓One platform spans GPU rental, RL training, evaluation, sandboxes and inference
- ✓Spot instances discounted up to 90% versus on-demand
- ✓Stack proven on the company's own 106B INTELLECT-3 run and by customers such as Ramp and Zapier
- ✓Well capitalised after a $130M Series A in July 2026
Cons
- ✗No formal SLAs; reliability depends on the underlying third-party GPU providers
- ✗No published SOC 2, HIPAA or SSO documentation in the docs FAQ or on the homepage
- ✗Instance data is permanently deleted on termination or when credits run out, so persistence needs planning
- ✗Independent analysis (Implicator) notes INTELLECT-3 is a post-train of GLM-4.5-Air that trails GLM-4.6 on AIME, and calls the business model unproven
- ✗Multi-node clusters are documented only in US data centers
- ✗Lab training and dedicated inference prices are not published
Pricing
On-demand GPU instances
$2.43/GPU-hr (H100 on-demand, listed Sept 2026)
- ✓Non-interruptible
- ✓Credits deducted per minute while running
- ✓Rates vary by underlying provider and GPU type
Spot GPU instances
$0.94/GPU-hr (H100 spot, listed Sept 2026)
- ✓Up to 90% below on-demand
- ✓Interruptible capacity
Multi-node clusters
From ~$40.80/hr (sample H100 cluster, Ethernet); ~$52.80/hr with InfiniBand
- ✓H100 clusters in US data centers
- ✓Slurm or Kubernetes orchestration
Lab, inference and reserved clusters
Contact for pricing
- ✓Hosted RL training and evaluations
- ✓Per-token LoRA and serverless inference
- ✓Reserved clusters via sales call
GPU compute is usage-based: credits are deducted per minute while instances run, and the homepage listed H100 at $2.43/hr on-demand and $0.94/hr spot in September 2026, with rates varying by provider. Spot is up to 90% cheaper but interruptible, pods are deleted if credits run out, sample multi-node H100 clusters run about $40.80-$52.80/hr, and there are no formal SLAs. Lab training and dedicated inference list prices are not published; inference is per-token and larger commitments go through a sales call.
Security & Compliance
Connect
Sources
This page was written from 13 sources, 6 on domains other than primeintellect.ai.
- 1.primeintellect.ai — primeintellect.aivendor
- 2.primeintellect.ai — labvendor
- 3.primeintellect.ai — series avendor
- 4.primeintellect.ai — intellect 3vendor
- 5.primeintellect.ai — environmentsvendor
- 6.primeintellect.ai — fundraisevendor
- 7.primeintellect.ai — computevendor
- 8.docs.primeintellect.ai — faq
- 9.github.com — PrimeIntellect ai
- 10.github.com — prime rl
- 11.techcrunch.com — prime intellect raises 130m series a to help enterprises bui
- 12.implicator.ai — prime intellects intellect 3 open source ambition meets cent
- 13.en.wikipedia.org — Prime Intellect
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Gimlet Cloud
Multi-silicon inference cloud that splits AI agent workloads across GPUs, CPUs and accelerators
DigitalOcean Managed Agents
Managed agent runtime with microVM sandboxes, 16,000+ governed tools and serverless inference, billed on active CPU
Modular
MAX inference framework and Mojo language for serving AI models on NVIDIA, AMD and other chips
ZML/LLMD
Free, Python-free LLM inference server that runs open models on NVIDIA, AMD, Google TPU, Intel and Apple chips from one binary