Vector Core Compute
by Vector Core Compute (VC2)
Disaggregated enterprise inference cloud running CPUs, GPUs and RDUs in one pipeline
Vector Core Compute (VC2) is an enterprise inference cloud, launched by Vista Equity Partners and Cambium Capital, that splits AI inference across Intel Xeon CPUs, NVIDIA Blackwell GPUs and SambaNova RDUs instead of running everything on GPUs. It targets enterprises and AI platform teams running agentic workloads that need low-latency, US-metro-local inference capacity.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Contact for pricing
- Target Market
- CTOs, CIOs, VPs of Engineering, AI Infrastructure Leaders, Enterprise Architects
- Founded
- 2026
Key Features
- ✓Fully disaggregated inference
Each inference stage is assigned to a different processor type rather than running the whole pipeline on GPUs.
- ✓Tri-silicon architecture
Intel Xeon 6 CPUs for orchestration and tool execution, NVIDIA Blackwell GPUs for prefill and prompt caching, and SambaNova SN40 RDUs for decode.
- ✓Metro-distributed capacity
Inference endpoints are being distributed across 50-plus US metropolitan markets instead of concentrated in remote mega-sites, putting compute near the enterprises using it.
- ✓Agentic-workload orchestration
CPU-side orchestration and tool execution is designed specifically for multi-step agentic AI rather than single-shot prompts.
- ✓Independently benchmarked speed
Artificial Analysis measured the disaggregated architecture at least two to three times faster than a GPU-only stack.
Use Cases
- •Serve production agentic AI at scale
Run multi-step agent workloads where CPU-based tool execution and orchestration sit alongside GPU prefill and RDU decode in a single pipeline.
- •Low-latency regional inference
Place inference endpoints in the same metro as the applications and customers consuming them to cut round-trip latency.
- •Diversify away from GPU-only capacity
Enterprises constrained by GPU supply or GPU-only economics can buy inference capacity built on a mixed CPU/GPU/RDU fleet.
Ideal For
Best For
- ✓High-volume agentic AI inference that needs low latency close to end users
- ✓Enterprises seeking an alternative to GPU-only inference capacity
- ✓AI platform teams benchmarking cost-per-token against hyperscaler inference
Market Analysis
Pros
- ✓Independently benchmarked at two to three times faster than a GPU-only stack
- ✓Metro-distributed footprint reduces inference latency for US enterprises
- ✓Anchor customer Together AI and Vista's portfolio give it immediate demand
Cons
- ✗Only the Los Angeles site is live; Chicago, Seattle and Phoenix are still in development
- ✗No public pricing, published API documentation or self-service onboarding yet
- ✗Very new company with a short operating track record
Pricing
Enterprise inference capacity
Contact for pricing
- ✓Disaggregated CPU/GPU/RDU inference
- ✓Metro-local endpoints
- ✓Agentic workload orchestration
No public pricing has been published; capacity is sold to enterprise customers directly.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Infinity Ignition
AI research agent that writes and optimizes inference kernels to make any AI chip production-ready in days
DeepInfra
Purpose-built inference cloud serving 200+ open-source AI models on OpenAI-compatible APIs
Alibaba Cloud Agent Native Cloud
Agent-native cloud suite for building, running, governing and observing enterprise AI agents at scale
Couchbase AI Data Plane
The operational data layer for production AI agents — persistent agent memory, context retrieval, tools and traces in one platform