Arcee Trinity
by Arcee AI
US-built open-weight model family, from on-device Trinity Nano to the 400B-parameter Trinity Large, that you can run on your own infrastructure
Arcee Trinity is a family of open-weight language models from Arcee AI, spanning the on-device Trinity Nano, the general-purpose Trinity Mini and the 400-billion-parameter sparse mixture-of-experts Trinity Large. It is built for enterprises, developers and public-sector teams that want frontier-class reasoning and agent capability while keeping control of the weights, the hosting and their data.
Arcee Trinity is the open-weight model family developed by Arcee AI, a San Francisco company founded in 2023 by Mark McQuade, who previously worked at Hugging Face. Arcee originally sold post-training and fine-tuning infrastructure before committing most of its capital to training its own models. The lineup runs from Trinity Nano, small enough for on-device use, through Trinity Mini, an everyday workhorse, to Trinity Large, a sparse mixture-of-experts model with roughly 400 billion total parameters and 13 billion active per token (256 experts, four active). Trinity Large was pretrained on 17 trillion tokens over 33 days on 2,048 NVIDIA B300 GPUs, with data curation by DatologyAI, and Arcee says its entire 2025 model lineup cost about $20 million to build. It ships as Preview, Base and TrueBase checkpoints plus a reasoning variant, Trinity-Large-Thinking, and Arcee cites up to 512k-token context support. Weights are downloadable from Hugging Face, including FP8, NVFP4, W4A16 and GGUF quantisations, and serve on vLLM or SGLang; Arcee says Trinity is moving to the OpenMDW-1.1 licence. For teams that do not want to self-host, Arcee's API charges $0.25 per million input tokens and $0.80 per million output tokens for Trinity-Large-Thinking, which is also listed on OpenRouter, and the same platform serves third-party open models such as DeepSeek, GLM and Kimi. On September 16, 2026 Arcee announced a Series B led by Vista Equity Partners, Cambium Capital and Emergence Capital at a valuation above $1 billion, reportedly worth at least $150 million, and it is working with the US Department of Energy's national laboratories on Genesis-Science-1, an open model for scientific computing.
Platform and ML engineering leaders, especially in regulated, sovereign or public-sector settings, who need a capable US-built model they can self-host, inspect and fine-tune rather than rent from a closed API.
Frontier-scale open weights with full control over hosting and data, and a low-cost hosted API ($0.25/$0.80 per million tokens) when self-hosting is not worth it.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Free
- Target Market
- CTOs, Heads of AI, ML Engineers, Enterprise Developers
- Deployment
- Self-hosted, API-based
- Founded
- 2023
- Headquarters
- San Francisco, United States
Key Features
- ✓Trinity Large (400B sparse MoE)
Roughly 400B parameters with only 13B active per token, delivering large-model quality at a fraction of dense-model inference cost.
- ✓Trinity-Large-Thinking
Reasoning variant aimed at multi-step agentic tasks and tool calling, served on vLLM and SGLang with reasoning and tool-call parsers.
- ✓Base and TrueBase checkpoints
Raw pretraining checkpoints, including a 10T-token checkpoint without instruct data, released for teams that want to post-train themselves.
- ✓Trinity Mini and Trinity Nano
Smaller models for everyday workloads and on-device inference, so one model family covers edge devices through the data center.
- ✓Quantised releases
FP8, NVFP4, W4A16 and GGUF variants on Hugging Face reduce the hardware needed to self-host the models.
- ✓Hosted API and OpenRouter
Pay-per-token API access lets teams evaluate or run Trinity in production without operating their own GPU clusters.
Capabilities
Use Cases
- •Sovereign or private-cloud assistant
A regulated enterprise self-hosts Trinity Large inside its own cloud account so prompts and outputs never leave its perimeter.
- •Domain fine-tuning
An ML team post-trains the Trinity Base checkpoint on proprietary documents to build a specialised model that it fully owns.
- •Low-cost agent backend
Developers route agent and tool-calling workloads to Trinity-Large-Thinking through the API at $0.25 input and $0.80 output per million tokens.
- •Scientific research models
Arcee and the US Department of Energy's national laboratories are building Genesis-Science-1, an open-weight model for scientific computing and research tasks.
Ideal For
Best For
- ✓Organisations requiring self-hosted models for data sovereignty or private-cloud deployment
- ✓Public-sector and research teams that prefer US-developed open-weight models
- ✓Post-training a base model on proprietary data using the Base and TrueBase checkpoints
- ✓Cost-sensitive reasoning and tool-calling workloads through a low-price hosted API
- ✓On-device or edge inference with the small Trinity Nano model
Not Ideal For
- ✗Teams that need best-in-class coding-agent or agentic SQL performance today: a Hacker News practitioner found a much smaller Qwen model scored higher on an agentic SQL test, and Arcee itself flagged rough edges in coding agents.
- ✗Buyers who need a vendor-certified hosted service with published SOC 2, SLAs and data-processing terms; Arcee's site lists no security certifications.
- ✗Organisations without GPU infrastructure or MLOps staff that still want to self-host, since a 400B mixture-of-experts model needs multi-GPU serving expertise.
Deployment
Market Analysis
Pros
- ✓Open weights you can inspect, fine-tune and host anywhere
- ✓Very low hosted API price and an efficient sparse architecture
- ✓A complete family from on-device Nano to 400B Large
- ✓Well funded ($1B+ valuation) with a US Department of Energy collaboration
Cons
- ✗Mixed practitioner results: a Hacker News user scored Trinity-Large-Thinking 16-17/25 on an agentic SQL test versus 23/25 for a 27B Qwen model
- ✗The model card warns multi-turn agents must preserve reasoning content across turns or the model can produce malformed tool calls
- ✗The licence is changing (Arcee says Trinity is moving to OpenMDW-1.1), so legal teams must review terms per checkpoint
- ✗No published security certifications, SLAs or named customer references for the hosted API
Pricing
Open weights (self-hosted)
$0
- ✓Download from Hugging Face
- ✓Serve on vLLM or SGLang
- ✓Quantised FP8/NVFP4/GGUF variants
Arcee API: Trinity-Large-Thinking
$0.25 per 1M input tokens / $0.80 per 1M output tokens
- ✓$0.06 per 1M cached input tokens
- ✓Also available via OpenRouter
Enterprise
Contact for pricing
- ✓Sales-led engagements for customisation and deployment
The weights are free to download and self-host, so the real cost of self-hosting is your own multi-GPU infrastructure and operations. Arcee's hosted API is metered per token: Trinity-Large-Thinking costs $0.25 per million input tokens, $0.80 per million output tokens and $0.06 per million cached input tokens, matching OpenRouter's listing. Enterprise terms are sales-led and unpublished.
Security & Compliance
Sources
This page was written from 9 sources, 7 on domains other than arcee.ai.
- 1.arcee.ai — arcee.aivendor
- 2.arcee.ai — trinity largevendor
- 3.docs.arcee.ai — pricing.md
- 4.huggingface.co — Trinity Large Thinking
- 5.openrouter.ai — arcee ai
- 6.siliconangle.com — open weight model developer arcee ai reaches 1b plus valuati
- 7.fortune.com — arcee ai trained four models for 20 million now its worth 1
- 8.globenewswire.com — arcee ai reaches 1b valuation with series b funding to advan
- 9.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Deep Cogito Cogito v2.1
MIT-licensed 671B hybrid-reasoning open model with short reasoning chains, plus custom post-training on enterprise data
Abacus.AI Smaug
Open-weight Smaug Agentic, Flash and Mini models fine-tuned for long-running enterprise AI agents
GPT-6 Astra
OpenAI's frontier model for autonomous computer use, gated cyber capability and long-horizon coding
Inkling
Apache-2.0 975B-parameter open-weights model from Mira Murati's lab, built to be fine-tuned rather than rented