Databricks Mosaic AI
by Databricks
Build, evaluate, serve and govern AI agents on the data they already run on
Mosaic AI is the AI and agent layer of the Databricks platform, spanning agent authoring, evaluation, vector search, model serving and governance. It lets teams build agents against governed enterprise data in one environment with lineage and access control intact. It suits organisations already standardised on Databricks, and is heavy for those that are not.
Mosaic AI is the generative-AI and agent layer of the Databricks platform, assembled largely from the $1.3 billion MosaicML acquisition Databricks announced in June 2023 and since folded into the wider lakehouse stack. It is not one product but a set of interlocking components: Agent Framework and Agent Bricks for authoring agents fine-tuned on enterprise data using synthetic data, custom evaluation and automated tuning; Agent Evaluation, which grades output against golden examples using LLM judges and lets subject-matter experts review responses through a review app without needing Databricks accounts; AI Search, a vector database with real-time syncing of source data for RAG; Model Serving, a serverless CPU and GPU endpoint layer that hosts proprietary models from Azure OpenAI, Bedrock and Anthropic alongside open-source Llama and Mistral, custom fine-tunes and classical scikit-learn or PyFunc models behind one API; managed MLflow, whose tracing records every step of agent inference for debugging and evaluation-dataset construction; and Unity Catalog plus the AI Gateway — rebranded Unity AI Gateway — which apply permissions, rate limits, budgets, quality monitoring and lineage across every model and MCP service in the estate. Agents can be written with LangGraph, LangChain, OpenAI or LlamaIndex and deployed through Model Serving or Databricks Apps. Databricks reports over 20,000 customer organisations including 70% of the Fortune 500, and publishes agent outcomes from FactSet (44% accuracy improvement), Block ($10M in productivity gains), Intercontinental Exchange (96% response accuracy) and Comcast (10x cost reduction).
The data platform or AI leader at an organisation whose governed enterprise data already lives in Databricks and who needs agents built against it without exporting the data.
Agents built, evaluated, served and governed in the same environment as the data, with Unity Catalog permissions and lineage intact end to end.
At a Glance
- Category
- Enterprise Platform
- Pricing
- Usage-based, Contact for pricing
- Target Market
- CIOs, CTOs, Data Scientists, Enterprise Developers, MLOps Engineers
- Deployment
- Cloud-only, Multi-cloud
- Founded
- 2013
- Headquarters
- San Francisco, United States
- Team Size
- 500+
- Customers
- 20,000+ organisations, including 70% of the Fortune 500
Key Features
- ✓Agent Bricks and Agent Framework
Builds agents fine-tuned on enterprise data using synthetic data generation, custom evaluation and automated tuning to optimise both quality and cost.
- ✓Agent Evaluation with LLM judges
Grades agent output against golden examples on accuracy and helpfulness, and lets subject-matter experts review responses without Databricks accounts.
- ✓Model Serving
Serverless CPU and GPU endpoints host proprietary, open-source, custom fine-tuned and classical ML models behind one unified interface and API.
- ✓AI Search vector database
High-performance vector indexes with real-time syncing from source Delta tables, so RAG retrieval does not drift from the underlying data.
- ✓Unity AI Gateway
Applies governance, budgets, rate limits, permissions and threshold-based monitoring across every LLM and MCP service in the enterprise.
- ✓MLflow tracing
Records each step of model and agent inference to debug performance problems and build evaluation datasets from real production traffic.
- ✓Unity Catalog governance
Enforces access controls and tracks data and tool lineage across agent workflows, including the function and MCP tool registry.
Capabilities
Use Cases
- •Governed RAG over enterprise data
Index Delta tables in AI Search with real-time sync so retrieval stays current and Unity Catalog permissions still apply to every lookup.
- •Text-to-code knowledge agents
FactSet reports a 44% accuracy improvement on a text-to-code knowledge agent built on the platform against proprietary financial data.
- •Internal agent automation at scale
Block reports $10 million in productivity gains from AI agent automation running on the Databricks agent stack.
- •Consolidating multi-provider model access
Route Azure OpenAI, Bedrock, Anthropic and self-hosted open models through one gateway with shared budgets, rate limits and audit lineage.
- •Serving classical ML and GenAI together
Deploy scikit-learn and PyFunc models alongside fine-tuned LLMs on the same serverless endpoints and monitoring surface.
Ideal For
Best For
- ✓Existing Databricks customers extending governed lakehouse data into RAG and agent applications without moving it
- ✓Regulated enterprises that need model, agent and MCP access controlled by one permission and lineage layer
- ✓Teams running classical ML and generative AI side by side who want a single serving and monitoring surface
- ✓Organisations that need agent evaluation with human subject-matter-expert review built into the workflow
- ✓Multi-model estates mixing Azure OpenAI, Bedrock, Anthropic, Llama and custom fine-tunes behind one governed API
Not Ideal For
- ✗Teams that only need a hosted model API or a simple chatbot — reviewers consistently call the platform excessive for a plain inference endpoint, and a narrower product ships faster
- ✗Organisations not already on Databricks: the stack is coupled to Spark-based workflows and carries a steep learning curve for anyone who is not a data engineer
- ✗Buyers who need predictable, published unit costs — everything is DBU-metered and varies by cloud, region, compute type, model and endpoint capacity
- ✗Teams that need portability across clouds or onto custom infrastructure, where the platform's flexibility is limited
Integrations
Deployment
Market & Ratings
20,000+ organisations, including 70% of the Fortune 500
Market Analysis
Pros
- ✓Data, agent development, serving and governance sit in one environment, which removes the export-and-re-govern step most GenAI stacks require
- ✓Genuinely broad model coverage — Azure OpenAI, Bedrock, Anthropic, Llama, Mistral, custom fine-tunes and classical ML behind a single API
- ✓MLflow tracing gives visibility into each agent step rather than only the final output, and feeds evaluation datasets directly
- ✓Published customer outcomes are specific and attributable: FactSet 44% accuracy improvement, Block $10M in productivity gains, ICE 96% response accuracy, Comcast 10x cost reduction
- ✓Scale and durability are not in question at 20,000+ customer organisations and 70% of the Fortune 500
Cons
- ✗Heavy coupling to Spark-based workflows and an architecture reviewers describe as analytics-first rather than GenAI-first, which adds weight to simple application-side AI routing
- ✗Steep learning curve for anyone who is not a data engineer; the platform assumes Databricks, data engineering and ML expertise in-house
- ✗Consumption-based DBU pricing is opaque and hard to forecast — cost varies by cloud, region, compute type, model, tokens and endpoint capacity, and per-token DBU charges can distort unit economics at production scale
- ✗Limited flexibility deploying GenAI systems across clouds or onto custom infrastructure
- ✗Persistent branding churn makes the product hard to pin down: the Mosaic AI Gateway is now the Unity AI Gateway, and current Databricks agent documentation no longer uses the 'Mosaic AI Agent Framework' name at all
- ✗Reserved-capacity commitment billing converts variable spend into fixed cost, which cuts against the pay-as-you-go pitch
Pricing
Foundation Model APIs — pay per token
Usage-based, DBU per 1M tokens
- ✓DeepSeek V4 Flash at 2.000 input / 4.000 output / 0.400 cache-read DBU per 1M tokens
- ✓Llama 3.3 70B at 7.143 input / 21.429 output
- ✓GLM-5.2 and 5.3 at 20.000 / 62.857 / 3.714
- ✓Kimi K3 at 42.857 / 214.286 / 4.286
- ✓No capacity commitment
Provisioned throughput
Usage-based, DBU per hour
- ✓Qwen 3.5 122B at 85.714 DBU/hour on demand
- ✓GLM-5.2 at 142.857 DBU/hour on demand, 121.429 with a 1-month reservation
- ✓Batch inference from 20.000 to 85.714 DBU/hour
- ✓Guaranteed throughput
GPU Model Serving
Usage-based, DBU per hour
- ✓T4 at 10.48 DBU/hour
- ✓A10G single GPU at 20.00
- ✓L40S single GPU at 44.86
- ✓H100 single GPU at 100.00
- ✓H100 x8 at 800.00
- ✓14-day free trial
Committed use / Enterprise
Contact for pricing
- ✓Committed-use discounts
- ✓Custom requirements
- ✓Reserved capacity
Everything is metered in DBUs and there is no flat monthly fee, so the dollar cost depends on cloud, region, workspace agreement, compute type, model, tokens and endpoint capacity. Published rates are DBU-denominated: GPU serving runs 10.48 DBU/hour for a T4 up to 800.00 for an eight-way H100, pay-per-token ranges from 2.000 DBU per million input tokens on DeepSeek V4 Flash to 42.857 on Kimi K3, and provisioned throughput starts around 85.714 DBU/hour. Agent workloads also pull in separate data-processing costs. A 14-day free trial exists, committed-use discounts are quote-only, and reviewers flag that reserved-capacity commitments convert variable spend into fixed cost.
Security & Compliance
Connect
Sources
This page was written from 9 sources, 3 on domains other than databricks.com.
- 1.databricks.com — artificial intelligencevendor
- 2.databricks.com — model servingvendor
- 3.databricks.com — foundation model servingvendor
- 4.databricks.com — model servingvendor
- 5.databricks.com — about usvendor
- 6.databricks.com — compliancevendor
- 7.truefoundry.com — 8 best databricks mosaic ai alternatives
- 8.therundown.ai — databricks mosaic ai
- 9.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Pigment
AI-native integrated business planning with Modeler and Analyst agents, AI-generated planning apps and an MCP server
Red Hat AI
Run any model on any accelerator across hybrid cloud, with safety evidence and cost attribution built in
Genesys Cloud
Agentic orchestration for customer experience — AI agents, employees and systems on one governed platform
VMware Tanzu Platform
Production AI agents inside your own private cloud, with a deny-by-default runtime