D

Databricks Mosaic AI

by Databricks

Enterprise PlatformAI Agents & OrchestrationData & AnalyticsInfrastructure & Cloud

Build, evaluate, serve and govern AI agents on the data they already run on

Usage-based · Contact for pricing·Added Jun 21, 2026·Updated Sep 8, 2026
Share:
THE DAILY BRIEF
Databricks Mosaic AI

by Databricks

Enterprise PlatformAI Agents & OrchestrationData & AnalyticsInfrastructure & Cloud

Build, evaluate, serve and govern AI agents on the data they already run on

Usage-based · Contact for pricing

Mosaic AI is the AI and agent layer of the Databricks platform, spanning agent authoring, evaluation, vector search, model serving and governance. It lets teams build agents against governed enterprise data in one environment with lineage and access control intact. It suits organisations already standardised on Databricks, and is heavy for those that are not.

At a Glance

Category
Enterprise Platform
Pricing
Usage-based, Contact for pricing
Target Market
CIOs, CTOs, Data Scientists, Enterprise Developers, MLOps Engineers
Deployment
Cloud-only, Multi-cloud
Founded
2013
Headquarters
San Francisco, United States
Team Size
500+
Customers
20,000+ organisations, including 70% of the Fortune 500

Key Features

  • Agent Bricks and Agent Framework
  • Agent Evaluation with LLM judges
  • Model Serving
  • AI Search vector database
  • Unity AI Gateway
  • MLflow tracing
  • Unity Catalog governance

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Governed RAG over enterprise data
  • Text-to-code knowledge agents
  • Internal agent automation at scale
  • Consolidating multi-provider model access
  • Serving classical ML and GenAI together

Ideal For

Best For

  • Existing Databricks customers extending governed lakehouse data into RAG and agent applications without moving it
  • Regulated enterprises that need model, agent and MCP access controlled by one permission and lineage layer
  • Teams running classical ML and generative AI side by side who want a single serving and monitoring surface
  • Organisations that need agent evaluation with human subject-matter-expert review built into the workflow
  • Multi-model estates mixing Azure OpenAI, Bedrock, Anthropic, Llama and custom fine-tunes behind one governed API

Not Ideal For

  • Teams that only need a hosted model API or a simple chatbot — reviewers consistently call the platform excessive for a plain inference endpoint, and a narrower product ships faster
  • Organisations not already on Databricks: the stack is coupled to Spark-based workflows and carries a steep learning curve for anyone who is not a data engineer
  • Buyers who need predictable, published unit costs — everything is DBU-metered and varies by cloud, region, compute type, model and endpoint capacity
  • Teams that need portability across clouds or onto custom infrastructure, where the platform's flexibility is limited

Market Analysis

Enterprise-gradeGovernedData-platform-native

Pros

  • Data, agent development, serving and governance sit in one environment, which removes the export-and-re-govern step most GenAI stacks require
  • Genuinely broad model coverage — Azure OpenAI, Bedrock, Anthropic, Llama, Mistral, custom fine-tunes and classical ML behind a single API
  • MLflow tracing gives visibility into each agent step rather than only the final output, and feeds evaluation datasets directly
  • Published customer outcomes are specific and attributable: FactSet 44% accuracy improvement, Block $10M in productivity gains, ICE 96% response accuracy, Comcast 10x cost reduction
  • Scale and durability are not in question at 20,000+ customer organisations and 70% of the Fortune 500

Cons

  • Heavy coupling to Spark-based workflows and an architecture reviewers describe as analytics-first rather than GenAI-first, which adds weight to simple application-side AI routing
  • Steep learning curve for anyone who is not a data engineer; the platform assumes Databricks, data engineering and ML expertise in-house
  • Consumption-based DBU pricing is opaque and hard to forecast — cost varies by cloud, region, compute type, model, tokens and endpoint capacity, and per-token DBU charges can distort unit economics at production scale
  • Limited flexibility deploying GenAI systems across clouds or onto custom infrastructure
  • Persistent branding churn makes the product hard to pin down: the Mosaic AI Gateway is now the Unity AI Gateway, and current Databricks agent documentation no longer uses the 'Mosaic AI Agent Framework' name at all
  • Reserved-capacity commitment billing converts variable spend into fixed cost, which cuts against the pay-as-you-go pitch

Pricing

Foundation Model APIs — pay per token

Usage-based, DBU per 1M tokens

  • DeepSeek V4 Flash at 2.000 input / 4.000 output / 0.400 cache-read DBU per 1M tokens
  • Llama 3.3 70B at 7.143 input / 21.429 output
  • GLM-5.2 and 5.3 at 20.000 / 62.857 / 3.714
  • Kimi K3 at 42.857 / 214.286 / 4.286
  • No capacity commitment

Provisioned throughput

Usage-based, DBU per hour

  • Qwen 3.5 122B at 85.714 DBU/hour on demand
  • GLM-5.2 at 142.857 DBU/hour on demand, 121.429 with a 1-month reservation
  • Batch inference from 20.000 to 85.714 DBU/hour
  • Guaranteed throughput

GPU Model Serving

Usage-based, DBU per hour

  • T4 at 10.48 DBU/hour
  • A10G single GPU at 20.00
  • L40S single GPU at 44.86
  • H100 single GPU at 100.00
  • H100 x8 at 800.00
  • 14-day free trial

Committed use / Enterprise

Contact for pricing

  • Committed-use discounts
  • Custom requirements
  • Reserved capacity

Everything is metered in DBUs and there is no flat monthly fee, so the dollar cost depends on cloud, region, workspace agreement, compute type, model, tokens and endpoint capacity. Published rates are DBU-denominated: GPU serving runs 10.48 DBU/hour for a T4 up to 800.00 for an eight-way H100, pay-per-token ranges from 2.000 DBU per million input tokens on DeepSeek V4 Flash to 42.857 on Kimi K3, and provisioned throughput starts around 85.714 DBU/hour. Agent workloads also pull in separate data-processing costs. A 14-day free trial exists, committed-use discounts are quote-only, and reviewers flag that reserved-capacity commitments convert variable spend into fixed cost.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Mosaic AI is the AI and agent layer of the Databricks platform, spanning agent authoring, evaluation, vector search, model serving and governance. It lets teams build agents against governed enterprise data in one environment with lineage and access control intact. It suits organisations already standardised on Databricks, and is heavy for those that are not.

Mosaic AI is the generative-AI and agent layer of the Databricks platform, assembled largely from the $1.3 billion MosaicML acquisition Databricks announced in June 2023 and since folded into the wider lakehouse stack. It is not one product but a set of interlocking components: Agent Framework and Agent Bricks for authoring agents fine-tuned on enterprise data using synthetic data, custom evaluation and automated tuning; Agent Evaluation, which grades output against golden examples using LLM judges and lets subject-matter experts review responses through a review app without needing Databricks accounts; AI Search, a vector database with real-time syncing of source data for RAG; Model Serving, a serverless CPU and GPU endpoint layer that hosts proprietary models from Azure OpenAI, Bedrock and Anthropic alongside open-source Llama and Mistral, custom fine-tunes and classical scikit-learn or PyFunc models behind one API; managed MLflow, whose tracing records every step of agent inference for debugging and evaluation-dataset construction; and Unity Catalog plus the AI Gateway — rebranded Unity AI Gateway — which apply permissions, rate limits, budgets, quality monitoring and lineage across every model and MCP service in the estate. Agents can be written with LangGraph, LangChain, OpenAI or LlamaIndex and deployed through Model Serving or Databricks Apps. Databricks reports over 20,000 customer organisations including 70% of the Fortune 500, and publishes agent outcomes from FactSet (44% accuracy improvement), Block ($10M in productivity gains), Intercontinental Exchange (96% response accuracy) and Comcast (10x cost reduction).

Ideal Buyer

The data platform or AI leader at an organisation whose governed enterprise data already lives in Databricks and who needs agents built against it without exporting the data.

Key Benefit

Agents built, evaluated, served and governed in the same environment as the data, with Unity Catalog permissions and lineage intact end to end.

At a Glance

Category
Enterprise Platform
Pricing
Usage-based, Contact for pricing
Target Market
CIOs, CTOs, Data Scientists, Enterprise Developers, MLOps Engineers
Deployment
Cloud-only, Multi-cloud
Founded
2013
Headquarters
San Francisco, United States
Team Size
500+
Customers
20,000+ organisations, including 70% of the Fortune 500

Key Features

  • Agent Bricks and Agent Framework

    Builds agents fine-tuned on enterprise data using synthetic data generation, custom evaluation and automated tuning to optimise both quality and cost.

  • Agent Evaluation with LLM judges

    Grades agent output against golden examples on accuracy and helpfulness, and lets subject-matter experts review responses without Databricks accounts.

  • Model Serving

    Serverless CPU and GPU endpoints host proprietary, open-source, custom fine-tuned and classical ML models behind one unified interface and API.

  • AI Search vector database

    High-performance vector indexes with real-time syncing from source Delta tables, so RAG retrieval does not drift from the underlying data.

  • Unity AI Gateway

    Applies governance, budgets, rate limits, permissions and threshold-based monitoring across every LLM and MCP service in the enterprise.

  • MLflow tracing

    Records each step of model and agent inference to debug performance problems and build evaluation datasets from real production traffic.

  • Unity Catalog governance

    Enforces access controls and tracks data and tool lineage across agent workflows, including the function and MCP tool registry.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Governed RAG over enterprise data

    Index Delta tables in AI Search with real-time sync so retrieval stays current and Unity Catalog permissions still apply to every lookup.

  • Text-to-code knowledge agents

    FactSet reports a 44% accuracy improvement on a text-to-code knowledge agent built on the platform against proprietary financial data.

  • Internal agent automation at scale

    Block reports $10 million in productivity gains from AI agent automation running on the Databricks agent stack.

  • Consolidating multi-provider model access

    Route Azure OpenAI, Bedrock, Anthropic and self-hosted open models through one gateway with shared budgets, rate limits and audit lineage.

  • Serving classical ML and GenAI together

    Deploy scikit-learn and PyFunc models alongside fine-tuned LLMs on the same serverless endpoints and monitoring surface.

Ideal For

Best For

  • Existing Databricks customers extending governed lakehouse data into RAG and agent applications without moving it
  • Regulated enterprises that need model, agent and MCP access controlled by one permission and lineage layer
  • Teams running classical ML and generative AI side by side who want a single serving and monitoring surface
  • Organisations that need agent evaluation with human subject-matter-expert review built into the workflow
  • Multi-model estates mixing Azure OpenAI, Bedrock, Anthropic, Llama and custom fine-tunes behind one governed API

Not Ideal For

  • Teams that only need a hosted model API or a simple chatbot — reviewers consistently call the platform excessive for a plain inference endpoint, and a narrower product ships faster
  • Organisations not already on Databricks: the stack is coupled to Spark-based workflows and carries a steep learning curve for anyone who is not a data engineer
  • Buyers who need predictable, published unit costs — everything is DBU-metered and varies by cloud, region, compute type, model and endpoint capacity
  • Teams that need portability across clouds or onto custom infrastructure, where the platform's flexibility is limited

Integrations

SDK Available
SDK:PythonSQLScalaJavaR

Deployment

On-Premise

Market & Ratings

Estimated Customers

20,000+ organisations, including 70% of the Fortune 500

Market Analysis

Enterprise-gradeGovernedData-platform-native

Pros

  • Data, agent development, serving and governance sit in one environment, which removes the export-and-re-govern step most GenAI stacks require
  • Genuinely broad model coverage — Azure OpenAI, Bedrock, Anthropic, Llama, Mistral, custom fine-tunes and classical ML behind a single API
  • MLflow tracing gives visibility into each agent step rather than only the final output, and feeds evaluation datasets directly
  • Published customer outcomes are specific and attributable: FactSet 44% accuracy improvement, Block $10M in productivity gains, ICE 96% response accuracy, Comcast 10x cost reduction
  • Scale and durability are not in question at 20,000+ customer organisations and 70% of the Fortune 500

Cons

  • Heavy coupling to Spark-based workflows and an architecture reviewers describe as analytics-first rather than GenAI-first, which adds weight to simple application-side AI routing
  • Steep learning curve for anyone who is not a data engineer; the platform assumes Databricks, data engineering and ML expertise in-house
  • Consumption-based DBU pricing is opaque and hard to forecast — cost varies by cloud, region, compute type, model, tokens and endpoint capacity, and per-token DBU charges can distort unit economics at production scale
  • Limited flexibility deploying GenAI systems across clouds or onto custom infrastructure
  • Persistent branding churn makes the product hard to pin down: the Mosaic AI Gateway is now the Unity AI Gateway, and current Databricks agent documentation no longer uses the 'Mosaic AI Agent Framework' name at all
  • Reserved-capacity commitment billing converts variable spend into fixed cost, which cuts against the pay-as-you-go pitch

Pricing

Free Trial Available

Foundation Model APIs — pay per token

Usage-based, DBU per 1M tokens

  • DeepSeek V4 Flash at 2.000 input / 4.000 output / 0.400 cache-read DBU per 1M tokens
  • Llama 3.3 70B at 7.143 input / 21.429 output
  • GLM-5.2 and 5.3 at 20.000 / 62.857 / 3.714
  • Kimi K3 at 42.857 / 214.286 / 4.286
  • No capacity commitment

Provisioned throughput

Usage-based, DBU per hour

  • Qwen 3.5 122B at 85.714 DBU/hour on demand
  • GLM-5.2 at 142.857 DBU/hour on demand, 121.429 with a 1-month reservation
  • Batch inference from 20.000 to 85.714 DBU/hour
  • Guaranteed throughput

GPU Model Serving

Usage-based, DBU per hour

  • T4 at 10.48 DBU/hour
  • A10G single GPU at 20.00
  • L40S single GPU at 44.86
  • H100 single GPU at 100.00
  • H100 x8 at 800.00
  • 14-day free trial

Committed use / Enterprise

Contact for pricing

  • Committed-use discounts
  • Custom requirements
  • Reserved capacity

Everything is metered in DBUs and there is no flat monthly fee, so the dollar cost depends on cloud, region, workspace agreement, compute type, model, tokens and endpoint capacity. Published rates are DBU-denominated: GPU serving runs 10.48 DBU/hour for a T4 up to 800.00 for an eight-way H100, pay-per-token ranges from 2.000 DBU per million input tokens on DeepSeek V4 Flash to 42.857 on Kimi K3, and provisioned throughput starts around 85.714 DBU/hour. Agent workloads also pull in separate data-processing costs. A 14-day free trial exists, committed-use discounts are quote-only, and reviewers flag that reserved-capacity commitments convert variable spend into fixed cost.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Connect

Sources

This page was written from 9 sources, 3 on domains other than databricks.com.

  1. 1.databricks.comartificial intelligencevendor
  2. 2.databricks.commodel servingvendor
  3. 3.databricks.comfoundation model servingvendor
  4. 4.databricks.commodel servingvendor
  5. 5.databricks.comabout usvendor
  6. 6.databricks.comcompliancevendor
  7. 7.truefoundry.com8 best databricks mosaic ai alternatives
  8. 8.therundown.aidatabricks mosaic ai
  9. 9.hn.algolia.comhn.algolia.com
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe