Infinity Ignition
by Infinity Inc.
AI research agent that writes and optimizes inference kernels to make any AI chip production-ready in days
Infinity Ignition is an autonomous AI research agent that generates, tests, debugs and optimizes the low-level compute kernels determining how efficiently a chip runs AI models, compressing inference software development from years into days. It is aimed at silicon vendors, cloud providers and infrastructure teams deploying models on non-NVIDIA accelerators.
Infinity Inc., founded in 2025 by former Google Brain researcher and AGI House creator Jeremy Nixon, builds the software layer that makes any AI chip inference-ready. Its product, Ignition, is an autonomous AI research agent that writes the low-level code required to run models on NVIDIA-alternative silicon — generating debuggers, profilers, compilers, kernels and the orchestration of kernels in inference — then tests, debugs and benchmarks the hardware against that code and automatically rewrites it to improve throughput. The system builds decompilers to translate executable code back into manipulable source, constructs bit-for-bit accuracy checks and memory detection tooling, and self-improves by recording successes and building reusable problem representations. Reported results include a 34% inference throughput increase for Qwen3-8B versus the vLLM framework achieved in a single day, 92% of theoretical peak performance on d-Matrix's Corsair accelerator reached in 10 hours, and full on-chip deployment of Qwen3, Qwen3.5 and Gemma4 within 10 days. Infinity charges no upfront license fees; it takes a percentage cut of the performance gains and cost savings it delivers, measured in tokens per second — roughly 20% of the customer's compute purchase savings from throughput improvement. d-Matrix Corp is a confirmed customer, with discussions underway with other major chip and cloud companies, and the company says it is already generating millions of dollars in annual recurring revenue from chip design partnerships. On 20 July 2026 Infinity announced a $15 million seed round at a $100 million post-money valuation led by Touring Capital with Principal Venture Partners, plus angel participation from executives and researchers at OpenAI, Anthropic and chipmakers. The company had 26 employees at the time of the raise and will use the funding to scale Ignition, expand engineering and accelerate chip partnerships.
At a Glance
- Category
- Infrastructure & Cloud
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CTOs, Heads of AI Infrastructure, Silicon Vendors, Platform Engineering Leaders, ML Systems Engineers
- Founded
- 2025
Key Features
- ✓Autonomous kernel generation
Ignition writes the low-level compute kernels that determine how efficiently a chip runs AI models, without manual engineering.
- ✓Automated test, debug and rewrite loop
The agent measures hardware performance against generated code and automatically rewrites it when results fall short.
- ✓Full toolchain synthesis
Generates debuggers, profilers, compilers, kernels and the orchestration of kernels used in inference.
- ✓Bit-for-bit correctness verification
Builds decompilers and bit-for-bit accuracy checking plus memory detection tooling to validate generated code.
- ✓Self-improving research loop
Records successful optimizations and constructs reusable problem representations that improve subsequent runs.
- ✓Outcome-based commercial model
No upfront license fee — Infinity takes roughly 20% of the compute cost savings delivered, measured in tokens per second.
Capabilities
Use Cases
- •New accelerator bring-up
Reached 92% of theoretical peak performance on d-Matrix's Corsair accelerator within 10 hours.
- •Throughput optimization on existing hardware
Delivered a 34% inference throughput increase for Qwen3-8B versus the vLLM framework in a single day.
- •Rapid model porting to new silicon
Deployed Qwen3, Qwen3.5 and Gemma4 fully on chip within 10 days.
Ideal For
Best For
- ✓Silicon vendors bringing a new AI accelerator to production-ready inference
- ✓Infrastructure teams reducing inference cost by running models on non-NVIDIA chips
- ✓Extracting higher throughput from existing accelerator fleets without manual kernel engineering
Deployment
Market Analysis
Pros
- ✓Directly attacks the software bottleneck that keeps non-NVIDIA silicon out of production
- ✓Outcome-based pricing removes upfront risk for buyers
- ✓Concrete published benchmarks against vLLM and on d-Matrix hardware
- ✓Already generating millions in ARR from chip design partnerships
Cons
- ✗Very early stage — founded 2025 with 26 employees at the time of its seed round
- ✗Only one publicly confirmed customer (d-Matrix)
- ✗Sells primarily to silicon vendors, so most enterprises benefit indirectly
- ✗No public product documentation, SDK or self-serve access
Pricing
Outcome-based partnership
Contact for pricing
- ✓No upfront license fee
- ✓Approximately 20% of compute purchase savings from throughput gains
- ✓Savings measured in tokens per second
- ✓Chip design partnership engagement
Infinity charges no upfront license fees and instead takes a percentage cut of the performance gains and cost savings it delivers, reported at roughly 20% of the customer's compute purchase savings, measured in tokens per second.
Sources
This page was written from 2 sources, 2 on domains other than infinity.inc.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
WEKA NeuralMesh
Microservices storage and memory fabric built for AI training and inference at exabyte scale
Tsuga
Bring-your-own-cloud observability that keeps telemetry, and its cost, inside your own AWS account
OpenObserve
Open-source, Rust-based observability on object storage — logs, metrics, traces and LLM telemetry in one binary
Coralogix
AI-native observability that queries logs, metrics and traces in place — index-free, in your own S3 bucket