Galileo
by Cisco (Splunk) — formerly Galileo Technologies, Inc.
Evaluation, observability and runtime guardrails for GenAI and multi-agent systems
Galileo is an AI evaluation and observability platform where offline evals become production guardrails for LLM and multi-agent applications. Teams trace agent runs, score them with custom and built-in evaluators, and block bad outputs at runtime. Cisco acquired it in 2026 and is folding it into Splunk Observability as Agent Observability.
Galileo is an evaluation, observability and guardrail platform for generative-AI and agentic applications, founded by Vikram Chatterji and Atin Sanyal with a team drawn from Google AI, Google Brain and Apple Siri. Its organising idea is that the same evaluator should run in development and in production: teams build datasets, run offline evals for correctness, relevance, tool use, safety and security, inspect traces through a framework-agnostic graph engine with timeline and conversation views, and then promote those scoring functions into runtime guardrails that gate agent actions. The economics of that come from Luna-2, Galileo's family of small language models purpose-built for evaluation, which the company says runs 10 to 20 metrics simultaneously at sub-200ms latency with 100% sampling for roughly 97% less than LLM-as-judge approaches. An Insights Engine performs automatic failure detection and root-cause analysis across multi-agent coordination and tool use. It integrates with CrewAI, LangGraph, the OpenAI Agents SDK, LlamaIndex and Amazon Strands over OpenTelemetry, with Python and TypeScript SDKs and a REST API. In July 2025 Galileo made the full observability, evaluation and guardrail stack available on a free tier, and it has raised over $68 million from Battery Ventures, Scale Venture Partners, Databricks Ventures, Citi Ventures and Hugging Face CEO Clement Delangue, naming HP, Twilio, Reddit and Comcast as customers. Cisco announced its intent to acquire the company on 9 April 2026 and closed in May 2026; the technology now anchors Splunk Agent Observability, adding evaluation, tokenomics cost tracking and sub-200ms guardrails to Splunk Observability Cloud.
The platform or AI engineering team shipping multi-agent systems to production that needs the same evaluation logic to run as a pre-deploy gate and as a runtime guardrail.
Continuous, affordable scoring of every agent step — Luna-2 runs 10-20 metrics at 100% sampling in under 200ms — instead of sampled LLM-as-judge spot checks.
At a Glance
- Category
- Governance & Security
- Pricing
- Freemium, Subscription, Usage-based, Contact for pricing
- Target Market
- CTOs, CISOs, Enterprise Developers, Data Scientists, MLOps Engineers
- Deployment
- Cloud-first, Hybrid, Self-hosted
Key Features
- ✓Luna-2 evaluation models
Small language models purpose-built for scoring, running 10-20 metrics simultaneously at sub-200ms latency and roughly 97% lower cost than LLM-as-judge.
- ✓Agent observability graph engine
Framework-agnostic tracing with timeline and conversation views that captures multi-agent coordination and tool calls step by step.
- ✓Insights Engine
Automatically detects failure modes, links errors back to the originating trace and prescribes fixes rather than only flagging a bad score.
- ✓Runtime guardrails
Intercepts prompts and outputs in under 200ms to block prompt injection, hallucinations, tool misuse and PII, PHI or PCI leakage.
- ✓Offline-to-production evaluators
The same custom or built-in evaluator used on a dataset in development becomes the production guardrail, removing dev/prod scoring drift.
- ✓Tokenomics cost tracking
Tracks token usage and spend by request, model, agent and workflow, and correlates cost against output quality.
- ✓Framework and OTel integrations
Works with CrewAI, LangGraph, OpenAI Agents SDK, LlamaIndex and Amazon Strands over OpenTelemetry via Python and TypeScript SDKs.
Capabilities
Use Cases
- •Gating an agent release
Run offline evaluations against a golden dataset and block deployment when correctness, tool-use or safety scores regress against the prior build.
- •Catching hallucinations before the user does
Score every production response with Luna-2 at full sampling and intercept unsupported claims inside the sub-200ms guardrail window.
- •Debugging a multi-agent failure
Trace the failing session through the graph engine, then use the Insights Engine to identify which agent or tool call caused the break.
- •Controlling agent spend
Break token cost down by request, model, agent and workflow to find which step is expensive and whether it improves output quality.
- •Preventing sensitive-data leakage
Enforce runtime rules that block PII, PHI and PCI from leaving an agent, which matters for regulated healthcare and financial workloads.
Ideal For
Best For
- ✓Teams operating multi-agent systems that need step-level tracing plus session-level metrics across a whole agent journey
- ✓Enterprises that need runtime guardrails against prompt injection, hallucination and PII/PHI/PCI leakage before output reaches a user
- ✓Regulated organisations requiring VPC or on-premises deployment of their evaluation stack
- ✓Existing Splunk Observability Cloud customers who want AI agent monitoring correlated with underlying infrastructure and GPU telemetry
- ✓Teams already instrumented with OpenTelemetry, CrewAI, LangGraph, LlamaIndex or the OpenAI Agents SDK
Not Ideal For
- ✗Small teams that want fully public, self-serve pricing at production volume — the Free and Pro tiers cap at 5,000 and 50,000 traces per month and anything beyond that requires a sales conversation
- ✗Buyers who need the load-bearing capabilities on a self-serve plan: real-time guardrails, SSO, RBAC at enterprise grade, VPC and on-premises deployment are all Enterprise-only
- ✗Teams that want a vendor-neutral, open-source-first evaluation stack — Langfuse, Arize Phoenix and Comet Opik are open source, Galileo is not
- ✗Organisations that would rather not take a dependency on a product mid-absorption into a much larger vendor's portfolio
Integrations
Deployment
Market Analysis
Pros
- ✓Full observability, evaluation and guardrail stack is usable on a genuinely free tier with unlimited users and unlimited custom evaluators
- ✓Luna-2 changes the unit economics of evaluation — 10-20 metrics at 100% sampling and sub-200ms, at a claimed 97% saving over LLM judges
- ✓Named enterprise customers (HP, Twilio, Reddit, Comcast) and partners (MongoDB, CrewAI, Elastic) give the adoption claim substance
- ✓Cisco ownership removes vendor-viability risk and adds correlation with Splunk's existing infrastructure and GPU telemetry
- ✓Contributed its Agent Control framework to open source under Apache 2.0
Cons
- ✗Absent from the major independent 2026 evaluation-platform round-ups — Arize's seven-platform guide and MarkTechPost's eight-platform comparison both omit it entirely, so like-for-like third-party benchmarking is scarce
- ✗Everything an enterprise actually buys it for — real-time guardrails, SSO, VPC and on-premises deployment, dedicated inference — is Enterprise-only and quote-only
- ✗Trace-based metering means costs scale with agent verbosity: a step-level-instrumented agent fleet exhausts the 50,000-trace Pro tier fast, with no published price above it
- ✗The product is mid-absorption into Splunk: galileo.ai/products now redirects to Splunk's Agent Observability page, and the standalone Galileo brand has largely disappeared from that page
- ✗Effectively no practitioner discussion on Hacker News, and the 'Galileo AI' name collides with an unrelated UI-design tool, which makes independent research unusually hard
Pricing
Free
$0
- ✓5,000 traces per month
- ✓Unlimited users
- ✓Unlimited custom evaluations
- ✓Agent observability and Insights Engine
Pro
From $100/mo (billed yearly)
- ✓50,000 traces per month
- ✓Scales with trace volume
- ✓Standard RBAC
- ✓Advanced analytics and insights
- ✓Slack support
Enterprise
Contact for pricing
- ✓Unlimited traces and custom rate limits
- ✓Hosted, VPC or on-premises deployment
- ✓Enterprise-grade RBAC and SSO
- ✓Real-time guardrails
- ✓Low-latency dedicated inference servers
- ✓24/7 Slack, email and phone support
- ✓Dedicated CSM and forward-deployed engineering
Metered on traces, not seats — users are unlimited on every tier. Free gives 5,000 traces a month with unlimited custom evaluations; Pro starts at $100/month billed yearly for 50,000 traces and scales with volume. Enterprise is quote-only and is where the capabilities enterprises actually buy for sit: real-time guardrails, SSO, hosted/VPC/on-premises deployment, dedicated low-latency inference servers and 24/7 support. Budget accordingly — a production agent fleet emitting a trace per step will clear 50,000 traces quickly, and there is no published price above the Pro band. Splunk separately offers a free Observability tier for up to 15 hosts.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 6 on domains other than galileo.ai.
- 1.galileo.ai — galileo.aivendor
- 2.galileo.ai — pricingvendor
- 3.networkworld.com — cisco to acquire galileo for ai observability
- 4.blogs.cisco.com — cisco announces the intent to acquire galileo
- 5.prnewswire.com — galileo announces free agent reliability platform 302508172
- 6.splunk.com — splunk observability galileo
- 7.splunk.com — agent observability
- 8.arize.com — llm and agent evaluation platforms
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
WitnessAI
Network-layer AI security and governance for employee AI use, models, apps and agents
Ascerta
Enterprise AI management: measure the ROI, cost and adoption of every AI initiative, agent and coding tool
Reco
AI agent security and SaaS security platform that discovers, governs and secures every agent, app and identity
Arcjet
Runtime security for AI agents: observe, enforce and audit agent tool calls, with prompt-injection, PII and abuse controls in code