Abacus.AI Smaug
by Abacus.AI
Open-weight Smaug Agentic, Flash and Mini models fine-tuned for long-running enterprise AI agents
Smaug is Abacus.AI's line of three open-weight language models: Smaug Agentic, Smaug Flash and Smaug Mini. They are fine-tuned from Kimi K3, DeepSeek V4 Flash and Qwen3.8 27B for long-running agentic and coding loops. The line is for enterprise AI teams that want stronger agent performance they can self-host in their own VPC, or call cheaply through Abacus.AI's RouteLLM API.
Abacus.AI launched the Smaug line on 10 September 2026: three open-weight models fine-tuned for enterprise agentic workloads and published on Hugging Face. Smaug is a fine-tuning technique that Abacus.AI says works on any open-source base model. It combines human-curated agentic traces with synthetic data and masks reasoning tokens from the training loss to preserve the base model's reasoning. Smaug Agentic is built on Moonshot AI's Kimi K3: a 2.8-trillion-parameter mixture-of-experts model with 104 billion active parameters, a 1,048,576-token context window and a vision encoder, released under the inherited Kimi K3 License. It runs on vLLM, SGLang and TokenSpeed as a drop-in for Kimi K3 serving stacks and was evaluated on 8 B300 GPUs. It scores 69.9 on DeepSWE against 67.5 for the base, 64.6 against 62.2 on LiveBench agentic coding, and 94.1 against 93.5 on GPQA Diamond. Abacus.AI says it sustained a median of 78 agent steps across 113 DeepSWE tasks over seven hours with zero infrastructure errors. Smaug Flash is a 304B fine-tune of DeepSeek-V4-Flash-0731 with 1M context, released under MIT. It lifts LiveBench agentic coding from 46.8 to 61.1 and AutomationBench from 25.1 to 38.83, is aimed at personal agents on WhatsApp, Telegram and Slack, and costs $0.10 per million input tokens and $0.40 per million output tokens on RouteLLM. Smaug Mini is a 27B Apache 2.0 fine-tune of Qwen3.8-27B with 262K context, for multimodal chatbots and smaller reasoning jobs; it lifts JobBench from 33.4 to 50.5. Abacus.AI claims 15-20% better performance on long agent loops at no extra cost, and prices 10-100 times below Anthropic and OpenAI frontier models. Abacus.AI is based in San Francisco, founded by Bindu Reddy and Arvind Sundararajan, and backed by Coatue, Tiger Global, Index Ventures and Khosla Ventures. It reports more than 3 million users. It previously used the Smaug name for Smaug-72B in 2024.
The AI platform lead running high-volume coding or automation agents who wants to cut per-token spend against closed frontier models by self-hosting, or using a cheap API for, an open model tuned for agent loops.
Vendor-reported gains on agentic coding and automation benchmarks over the same open base models, at the same serving cost, deployable inside their own VPC.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Free, Subscription, Usage-based
- Target Market
- CTOs, Heads of AI, ML Platform Engineers, Enterprise Developers
- Deployment
- Open-source, Self-hosted, API-based
- Headquarters
- San Francisco, United States
- Customers
- Abacus.AI reports 3M+ users across professionals, small businesses and enterprises, including dozens of the Fortune 500 (company-wide, not Smaug-specific)
Key Features
- ✓Smaug Agentic (Kimi K3 base)
A 2.8T-parameter MoE model with 104B active parameters and 1M context, tuned for complex long-running coding loops. Abacus.AI positions it as a self-hostable replacement for Opus-class closed models.
- ✓Smaug Flash (DeepSeek V4 Flash base)
A 304B MIT-licensed model with large gains on agentic coding and AutomationBench over its base. It is priced at $0.10/$0.40 per million tokens on RouteLLM.
- ✓Smaug Mini (Qwen3.8 27B base)
A 27B Apache 2.0 multimodal model with 262K context for chatbots and smaller reasoning tasks. It gives agentic gains in a size that is practical to self-host.
- ✓Agent-trace fine-tuning method
Combines human-curated agentic traces with synthetic data and masks reasoning tokens from the loss. The aim is to improve tool-using behaviour without degrading the base model's reasoning.
- ✓Drop-in serving compatibility
The models run on vLLM, SGLang and TokenSpeed using the same serving recipes as their base models, so existing inference infrastructure can switch with little change.
- ✓OpenAI-compatible RouteLLM API
Smaug Flash is served through RouteLLM's OpenAI-compatible endpoint alongside 160+ other models. Teams can compare it with closed models without new integration work.
Capabilities
Use Cases
- •Replacing a closed model in a coding agent
An engineering platform team swaps an Opus-class closed model for self-hosted Smaug Agentic to cut per-token cost on multi-hour autonomous coding runs.
- •Messaging-channel personal agents
A company builds WhatsApp and Slack assistants on Smaug Flash through RouteLLM, paying $0.10 per million input tokens for high-volume conversational agent traffic.
- •In-VPC regulated agent workloads
A financial-services firm hosts Smaug models inside its own cloud VPC, so proprietary data used by automation agents never reaches a third-party model provider.
- •Compact multimodal support bot
A support organisation deploys the 27B Apache 2.0 Smaug Mini with 262K context to handle image-plus-text customer queries on modest self-hosted GPU capacity.
Ideal For
Best For
- ✓Long-running autonomous coding agents that need million-token context and stable multi-hour execution
- ✓Personal and messaging-channel agents on WhatsApp, Telegram or Slack built on the lower-cost Smaug Flash
- ✓Multimodal enterprise chatbots and smaller reasoning workloads on the 27B Apache 2.0 Smaug Mini
- ✓Regulated enterprises that must keep agent inference inside their own cloud VPC or GPU cluster
- ✓Teams already serving Kimi K3, DeepSeek V4 Flash or Qwen3.8 that want a drop-in, agent-tuned upgrade
Not Ideal For
- ✗Organisations without Blackwell-class GPU capacity that want to self-host the large models. Smaug Agentic is a 2.8T-parameter model evaluated on 8 B300s, and Smaug Flash's recommended serving targets Blackwell/GB300 hardware.
- ✗Buyers who need independently validated benchmarks before shortlisting. All published Smaug results are Abacus.AI's own evaluations, some run on different harnesses.
- ✗Teams that want a managed API for every size today. Only Smaug Flash appeared in the RouteLLM catalogue at research time, and Hugging Face listed no third-party inference providers for the new models.
- ✗Legal teams that only approve OSI-style licences. Smaug Agentic inherits the custom Kimi K3 License rather than MIT or Apache 2.0, so it needs separate review.
Deployment
Market & Ratings
Abacus.AI reports 3M+ users across professionals, small businesses and enterprises, including dozens of the Fortune 500 (company-wide, not Smaug-specific)
Market Analysis
Pros
- ✓Large reported lifts on agentic tasks for Smaug Flash (LiveBench agentic coding 61.1 vs 46.8; AutomationBench 38.83 vs 25.1)
- ✓Million-token context on Smaug Agentic and Smaug Flash for long-horizon agent sessions
- ✓Permissive MIT and Apache 2.0 licences on Flash and Mini allow private, commercial self-hosting
- ✓Very low hosted price for Smaug Flash ($0.10/$0.40 per million tokens) via an OpenAI-compatible API
- ✓Drop-in compatibility with existing vLLM and SGLang deployments of the base models
Cons
- ✗All benchmarks are Abacus.AI's own; TechEdgeAI notes some evaluations use different harnesses, so production validation is essential
- ✗Gains are marginal on several benchmarks: GPQA Diamond +0.6 for Smaug Agentic and +0.2 for Smaug Mini
- ✗Self-hosting the large models needs top-end hardware (Agentic evaluated on 8 B300s; Flash recommended on Blackwell/GB300)
- ✗At launch Hugging Face listed no third-party inference providers, and only Smaug Flash was in the RouteLLM catalogue
- ✗Very early adoption signal: hundreds of downloads or fewer per model on Hugging Face in the first week
- ✗The API route runs through ChatLLM, whose credit system an independent pricing guide calls opaque, citing Reddit complaints about forced credit top-ups and slow support
Pricing
Open weights (Hugging Face)
$0
- ✓Smaug Flash under MIT
- ✓Smaug Mini under Apache 2.0
- ✓Smaug Agentic under the Kimi K3 License
- ✓Self-host in VPC or on-premise GPU clusters
RouteLLM API + ChatLLM Teams
From $10/mo
- ✓$7 for the first month
- ✓Required to obtain RouteLLM API credentials
- ✓Smaug Flash at $0.10 input / $0.40 output per 1M tokens
- ✓OpenAI-compatible API to 160+ models
Enterprise
Contact for pricing
- ✓Custom pricing
- ✓Full API access per third-party pricing guide
All three models can be downloaded free, so self-hosting costs are infrastructure costs: Smaug Agentic was evaluated on 8 B300 GPUs. The managed route requires a ChatLLM Teams subscription ($7 for the first month, then $10 a month) for RouteLLM credentials, plus per-token billing; Smaug Flash is $0.10 per million input and $0.40 per million output tokens. Smaug Mini and Smaug Agentic did not appear in the RouteLLM catalogue at research time. An independent pricing guide criticises the ChatLLM credit system as opaque.
Security & Compliance
Connect
Sources
This page was written from 12 sources, 10 on domains other than abacus.ai.
- 1.abacus.ai — smaugvendor
- 2.abacus.ai — aboutvendor
- 3.prnewswire.com — abacusai launches the smaug line of open weight models optim
- 4.unite.ai — abacus ai releases three open weight smaug models for agenti
- 5.techedgeai.com — abacus ai smaug targets cheaper enterprise ai agents
- 6.huggingface.co — Smaug Agentic
- 7.huggingface.co — Smaug Flash
- 8.huggingface.co — Smaug Mini
- 9.huggingface.co — abacusai
- 10.routellm-apis.abacus.ai — routellm-apis.abacus.ai
- 11.eesel.ai — abacus ai pricing
- 12.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Deep Cogito Cogito v2.1
MIT-licensed 671B hybrid-reasoning open model with short reasoning chains, plus custom post-training on enterprise data
GPT-6 Astra
OpenAI's frontier model for autonomous computer use, gated cyber capability and long-horizon coding
Inkling
Apache-2.0 975B-parameter open-weights model from Mira Murati's lab, built to be fine-tuned rather than rented
Fundamental NEXUS
A foundation model built for tables, not text — enterprise prediction without feature engineering