General Compute
by General Compute Inc.
ASIC-based inference cloud built for agents, where latency compounds
General Compute is an inference cloud that serves open-weight LLMs on purpose-built ASIC accelerators instead of GPUs, targeting agentic workloads where dozens of sequential model calls make time-to-first-token the real bottleneck. It exposes an OpenAI-compatible API, advertises sub-300ms TTFT and 1,000+ tokens per second, and is designed so an autonomous agent can sign up and provision its own compute.
General Compute is an inference cloud that runs open-weight large language models on purpose-built ASIC accelerators rather than general-purpose GPUs, aimed squarely at agentic workloads where latency compounds. Its central argument is that an autonomous agent makes dozens of sequential model calls per task, so time-to-first-token — not aggregate throughput — sets the ceiling on how quickly the agent finishes; the company advertises sub-300ms time-to-first-token, sustained throughput above 1,000 tokens per second, up to 7x the speed of GPU-based alternatives, and a 99.9% uptime SLA. Architecturally the platform disaggregates the prefill and decode stages of inference so each can be scaled independently against the workload mix, and the racks are air-cooled at lower power density than GPU equivalents, sited in hydroelectric-powered data centres. The API is OpenAI-compatible with Node.js and Python SDKs, tool calling and JSON mode, streaming semantics matching OpenAI's, and prebuilt integrations for OpenClaw, OpenCode and the Vercel AI SDK; unusually, the docs and signup flow are explicitly built so an autonomous agent can register and provision its own inference without a human in the loop. Four models are offered — MiniMax M2.7 (192k context, $0.28 in / $1.20 out per million tokens), DeepSeek V3.2 (32k, $0.25 / $0.38, reasoning), DeepSeek V3.1 (128k, $0.21 / $0.79, reasoning) and GPT-OSS 120B (128k, $0.21 / $0.79) — with dedicated reserved capacity, private fine-tuned checkpoints and additional regions available on request beyond the default us-west-2. Founded by Jason Goodison and Finn Puklowski, the platform opened to early partners on 18 April 2026, reached general availability on 15 May 2026, and placed third on Product Hunt for 22 May 2026 with 331 upvotes.
An engineering team running latency-sensitive agent or voice workloads on open-weight models, where the product is bottlenecked by per-call round-trip time rather than by model capability.
Cuts the per-call latency that compounds across an agent's dozens of sequential LLM calls, without rewriting anything — the API is OpenAI-compatible.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Enterprise Developers, AI Infrastructure Engineers, ML Platform Teams
- Deployment
- Cloud-only, API-based
- Headquarters
- United States
Key Features
- ✓ASIC-based serving
Runs on purpose-built inference accelerators rather than Nvidia GPUs, which is the basis of the advertised latency advantage over general-purpose silicon.
- ✓Prefill/decode disaggregation
Separates the two inference stages so each scales independently with the workload mix, instead of one bottleneck constraining both.
- ✓OpenAI-compatible API
Endpoints, parameters and streaming semantics match OpenAI's, so an existing SDK integration migrates with a base-URL and key change.
- ✓Agent self-signup
Docs and the registration flow are built for machine consumption, so an autonomous agent can provision its own inference capacity without a human.
- ✓Tool calling and JSON mode
Structured-output support so agents can act on model responses safely, alongside reasoning-capable models with visible thinking.
- ✓Dedicated capacity and private models
Reserved infrastructure and customer-supplied fine-tuned checkpoints with private model IDs, priced on a custom basis for enterprise workloads.
- ✓Hydroelectric-powered, air-cooled data centres
Lower power density than GPU racks and renewable supply, which matters for buyers with reporting obligations on AI energy use.
Capabilities
Use Cases
- •Speeding up a coding agent
An agent making dozens of sequential calls per task finishes materially faster when each round trip drops, since the latency compounds across the chain.
- •Real-time voice assistants
Sub-300ms time-to-first-token keeps a spoken interaction inside the window where a pause still reads as natural to the caller.
- •Cost reduction on open-weight inference
Teams already running DeepSeek or GPT-OSS elsewhere can compare per-million-token rates published openly, starting at $0.21 input.
- •Serving a private fine-tuned model
Customers deploy their own checkpoint onto dedicated accelerator capacity with a private model ID rather than sharing pooled inference.
- •Agent-provisioned compute
An autonomous system registers, obtains an API key and buys its own inference capacity, useful for self-scaling agent fleets.
Ideal For
Best For
- ✓Coding agents and other multi-step workflows where dozens of sequential LLM calls make round-trip latency the dominant cost
- ✓Real-time voice AI where time-to-first-token determines whether the interaction feels conversational
- ✓Teams already standardised on open-weight models (DeepSeek, GPT-OSS, MiniMax) looking for a faster serving layer
- ✓Migrating an existing OpenAI-SDK codebase to cheaper open-weight inference with minimal code change
- ✓Deploying a private fine-tuned checkpoint onto dedicated accelerator capacity
Not Ideal For
- ✗Teams standardised on frontier proprietary models — no GPT-5.x, Claude or Gemini is offered, so this is not a drop-in replacement for those workloads
- ✗EU or APAC buyers with data-residency requirements: the default and only advertised region is us-west-2, and anything else requires a custom dedicated deal
- ✗Long-context agent workloads on the reasoning model, since DeepSeek V3.2 is capped at 32k here — short for agents that accumulate long tool-call histories
- ✗Enterprise procurement that requires published vendor viability signals; funding, headcount and security certifications are all undisclosed
- ✗Training or fine-tuning workloads — this is an inference layer, and custom checkpoints must be trained elsewhere and brought in
Integrations
Deployment
Market Analysis
Pros
- ✓Correctly identifies a real and under-served problem: agent latency compounds across sequential calls, and most inference clouds are tuned for chatbot-style single-turn throughput
- ✓Fully published per-token pricing with context windows in the public docs, which is more transparency than most inference vendors offer
- ✓OpenAI-compatible API with Python and Node SDKs makes evaluation genuinely cheap — a base-URL swap rather than a migration project
- ✓Product Hunt reception was solid for an infrastructure product: #3 of the day on 22 May 2026 with 331 upvotes and substantive technical questions in the comments
Cons
- ✗Only four models at launch, all open-weight — no Claude, GPT-5.x or Gemini — so it cannot serve teams standardised on frontier proprietary models, which is most enterprises
- ✗DeepSeek V3.2, the reasoning model, is capped at 32k context here, which is short for exactly the long-running agent workloads the platform markets itself to
- ✗All performance claims (7x faster, 1,000+ tokens/sec, sub-300ms TTFT) are vendor-published with no third-party benchmark, and Product Hunt commenters raised unanswered questions about ASIC model constraints and KV cache handling
- ✗Effectively no independent validation exists: no G2, Capterra or TrustRadius listing, and a Hacker News search for the company returns nothing about this product at all
- ✗Funding, headcount and ownership are entirely undisclosed, and no security certifications or trust page are published — a real vendor-viability and procurement problem for an infrastructure dependency
- ✗Single default region (us-west-2) with no advertised EU or APAC residency without a custom deal, and fixed-function ASIC silicon carries more architectural obsolescence risk than GPUs as model designs shift
Pricing
Self-serve API
From $0.21/1M input tokens
- ✓MiniMax M2.7 — $0.28 in / $1.20 out per 1M, 192k context
- ✓DeepSeek V3.2 — $0.25 in / $0.38 out per 1M, 32k context
- ✓DeepSeek V3.1 — $0.21 in / $0.79 out per 1M, 128k context
- ✓GPT-OSS 120B — $0.21 in / $0.79 out per 1M, 128k context
- ✓Signup credit
Dedicated capacity
Contact for pricing
- ✓Reserved accelerator infrastructure
- ✓Contractual SLAs and escalation paths
- ✓Dedicated regional deployments beyond us-west-2
Bring your model
Contact for pricing
- ✓Private fine-tuned checkpoints
- ✓Private model IDs
- ✓Dedicated infrastructure
Self-serve inference is metered per million tokens with rates published openly in the docs, from $0.21 in / $0.79 out for DeepSeek V3.1 and GPT-OSS 120B up to $0.28 in / $1.20 out for MiniMax M2.7. Dedicated capacity, non-default regions and private fine-tuned checkpoints are all custom-priced and gated behind a sales conversation, as are contractual SLAs and escalation paths. Signup credit is advertised inconsistently — $100 on the pricing page, $5 in the docs' OpenClaw guide, and a one-off $200 for the Product Hunt launch — so confirm the current amount before budgeting an evaluation.
Security & Compliance
Sources
This page was written from 6 sources, 4 on domains other than generalcompute.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Cisco Cloud Control
One console where humans and AI agents run and defend the whole Cisco estate
Vector Core Compute
Disaggregated enterprise inference cloud running CPUs, GPUs and RDUs in one pipeline
DeepInfra
Purpose-built inference cloud serving 200+ open-source AI models on OpenAI-compatible APIs
Infinity Ignition
AI research agent that writes and optimizes inference kernels to make any AI chip production-ready in days