Red Hat AI
by Red Hat (IBM)
Run any model on any accelerator across hybrid cloud, with safety evidence and cost attribution built in
Red Hat AI is a hybrid-cloud platform for building, serving and governing enterprise AI workloads on Kubernetes. It bundles Red Hat AI Enterprise, OpenShift AI, RHEL AI and Red Hat AI Inference so platform teams can run any model on any accelerator on-premises, at the edge or across clouds, with pre-deployment safety benchmarking, multi-tenant GPU controls and per-user token metering.
Red Hat AI is Red Hat's umbrella enterprise AI platform, built on OpenShift and positioned around a single claim: any model, any agent, any hardware accelerator, anywhere. It comprises Red Hat AI Enterprise (a unified model and application lifecycle environment), Red Hat OpenShift AI (predictive and generative model lifecycle at scale on Kubernetes), Red Hat Enterprise Linux AI for single-server LLM deployment, and Red Hat AI Inference, which is powered by vLLM. Version 3.5, generally available 9 September 2026, moved the platform's centre of gravity from getting models into production to running them as governed infrastructure. Its headline addition is EvalHub, now GA: an evaluation orchestration service with a versioned REST API that runs multi-framework benchmarks inside isolated, horizontally scalable Kubernetes jobs, letting teams pre-screen custom models, RAG configurations and agents for risk and emit auditable compliance reports. The validated model catalog gained 20-plus entries including Google Gemma 4, NVIDIA Nemotron 3 and Alibaba Qwen, each carrying Garak security scores and PII-exposure and toxicity assessments. On infrastructure, 3.5 added fair-share GPU scheduling, priority-aware serving that protects real-time inference from background jobs, dynamic GPU capacity borrowing, and hosted control planes on OpenShift Virtualization for VM-level tenant isolation on shared hardware. The llm-d distributed inference scheduler went generally available on CoreWeave CKS and Microsoft Azure with Amazon EKS in technology preview — the first time it runs beyond OpenShift. For agents, the Responses API and AutoRAG reached GA with pgvector and multilingual document support, NVIDIA NeMo Guardrails intercept malicious tool calls, and AI Hub ships starter templates for code review, document processing and research. Observability added GPU utilisation and inference-health dashboards, MLflow visual agentic tracing, and per-user token metering for model-as-a-service showback.
The platform engineering or MLOps leader at an enterprise already standardised on OpenShift or Kubernetes, who must serve multiple internal AI tenants from shared GPUs and produce audit evidence for what those models and agents do.
One governed control plane for models, agents and GPUs that works identically on-premises, at the edge and on AWS, Azure or CoreWeave — with pre-deployment safety scores and per-user token showback as first-class outputs rather than bolt-ons.
At a Glance
- Category
- Enterprise Platform
- Pricing
- Subscription, Contact for pricing
- Target Market
- CIOs, CTOs, Platform Engineers, MLOps Teams, Enterprise Architects
- Deployment
- Hybrid, Self-hosted, Multi-cloud, Cloud-first
- Founded
- 1993
- Headquarters
- Raleigh, United States
- Team Size
- 500+
Key Features
- ✓EvalHub safety benchmarking
Generally available evaluation orchestration service that runs multi-framework benchmarks in isolated Kubernetes jobs and emits auditable compliance reports before deployment.
- ✓Validated model catalog
Over twenty vetted models including Gemma 4, Nemotron 3 and Qwen, each shipping Garak security scores plus PII-exposure and toxicity risk assessments.
- ✓llm-d distributed inference
Distributed inference scheduler now generally available on CoreWeave CKS and Microsoft Azure, with Amazon EKS in technology preview and new KV-cache reuse scorers.
- ✓Multi-tenant GPU controls
Fair-share scheduling, priority-aware serving and dynamic capacity borrowing let several tenants share GPUs without background jobs starving real-time inference.
- ✓Per-user token metering and observability
Dashboards for GPU utilisation, inference health and model performance, plus token metering that makes model-as-a-service chargeback and showback possible.
- ✓AutoRAG and agent tooling
GA Responses API, pgvector-backed multilingual RAG, NeMo Guardrails that intercept malicious tool calls, and starter templates for code review and document workflows.
- ✓Hosted control planes on OpenShift Virtualization
Gives each tenant a dedicated cluster control plane with VM-level isolation while still sharing the underlying accelerator hardware.
Capabilities
Use Cases
- •GPU-as-a-service for internal teams
A platform team sells shared accelerator capacity internally, using fair-share scheduling to prevent contention and token metering to attribute cost per department.
- •Regulated model approval
A risk function runs EvalHub against a fine-tuned model and its RAG pipeline, producing safety and PII-exposure evidence before the model reaches production.
- •Sovereign or on-premises inference
An organisation that cannot send prompts to a public API serves open models in its own datacentre while keeping the same tooling as its cloud footprint.
- •Multi-cloud inference standardisation
A team runs llm-d on CoreWeave and Azure alongside OpenShift on-premises, keeping one scheduler and one operational model across all three environments.
- •Governed internal agents
Developers build document-processing or code-review agents from AI Hub templates while NeMo Guardrails block malicious tool calls and MLflow traces every agent step.
Ideal For
Best For
- ✓Enterprises already running OpenShift that want AI serving on the same control plane as their existing workloads
- ✓Regulated organisations needing auditable pre-deployment safety and compliance evidence for custom models, RAG systems and agents
- ✓Platform teams chargebacking shared GPU capacity across multiple internal tenants with fair-share scheduling and per-user token metering
- ✓Hybrid and sovereign deployments where inference must stay on-premises, at the edge or in a specific jurisdiction rather than in a vendor's cloud
- ✓Teams standardising on open inference components — vLLM and llm-d — to avoid lock-in to a single cloud's model-serving stack
Not Ideal For
- ✗Small teams or startups that just want an API key: this is heavyweight Kubernetes platform software, and G2 reviewers of the predecessor OpenShift Data Science product consistently name setup and configuration complexity as the main drawback
- ✗Organisations with no Kubernetes or OpenShift skills in-house — OpenShift-specific concepts like Routes and Operators add a learning curve on top of AI tooling that is itself unfamiliar
- ✗Buyers who need transparent list pricing: Red Hat publishes subscription guides but not prices, and accelerator subscriptions are quoted per deployment
- ✗Anyone wanting a turnkey frontier model service; Red Hat serves open and third-party models but does not train a competitive frontier model of its own
Integrations
Deployment
Market Analysis
Pros
- ✓Genuinely hybrid — the same platform and tooling run on-premises, at the edge and on AWS, Azure and CoreWeave, which is rare among enterprise AI platforms and decisive for regulated or sovereign workloads
- ✓Governance is built in rather than bought separately: pre-deployment safety benchmarking, model risk scoring, guardrails on tool calls and agent tracing all ship in the platform
- ✓Cost attribution is solved at the platform layer with per-user token metering and showback dashboards, which is what makes internal GPU-as-a-service chargeback workable
- ✓Backed by Red Hat's subscription support model and IBM ownership, so a platform bet here does not carry the vendor-survival risk of a venture-stage alternative
- ✓Open inference foundation (vLLM, llm-d) reduces lock-in relative to a hyperscaler's proprietary serving stack
Cons
- ✗Operational complexity is the recurring theme in independent reviews: G2 reviewers of the predecessor Red Hat OpenShift Data Science product repeatedly cite a steep initial learning curve and difficult setup and configuration, and the broader OpenShift line is described as heavy compared with plain Kubernetes
- ✗It presumes an OpenShift investment — OpenShift-specific concepts such as Routes and Operators are extra surface area on top of the AI tooling, which makes this a poor fit for smaller teams or simple workloads
- ✗No published pricing at all, and the self-managed accelerator subscription meters per physical GPU, so cost grows with hardware rather than usage and is hard to model before a quote
- ✗Practitioner mindshare is thin relative to the hyperscalers: Hacker News threads about OpenShift AI are almost entirely low-engagement press-release reposts, and the most-discussed item was a 2025 root-access vulnerability allowing full cluster takeover rather than a capability
- ✗Several 3.5 headline items are not production-ready — Amazon EKS support for llm-d is technology preview, vLLM Omni multimodal serving is early access, and storage offloading and the Kubeflow Spark Operator are developer preview
Pricing
Red Hat AI Enterprise
Contact for pricing
- ✓Unified model and application lifecycle
- ✓Unified entitlement covering unlimited hardware accelerators per node
- ✓EvalHub safety benchmarking
- ✓Validated model catalog
Red Hat OpenShift AI (self-managed)
Contact for pricing
- ✓Per-AI-Accelerator subscription for each physical GPU, TPU, NPU, FPGA or DPU
- ✓Runs on customer-managed OpenShift
- ✓Hybrid cloud, on-premises and edge deployment
Red Hat Enterprise Linux AI
Contact for pricing
- ✓Single-server LLM deployment
- ✓Bootable RHEL image with inference runtime
Red Hat publishes subscription guides but no list prices, so every deal is quoted. The metering unit is what matters: self-managed OpenShift AI requires a separate Red Hat AI Accelerator subscription for each discrete physical accelerator card — GPU, TPU, NPU, FPGA or DPU — that is not part of the CPU package, at the same rate regardless of OpenShift edition, so cost scales with GPU count rather than users or tokens. Red Hat AI Enterprise instead provides a unified entitlement covering unlimited accelerators per node, which changes the arithmetic sharply for dense GPU servers. Trials are offered; there is no free tier.
Security & Compliance
Sources
This page was written from 6 sources, 4 on domains other than redhat.com.
- 1.redhat.com — aivendor
- 2.redhat.com — red hat puts safety and observability core enterprise ai redvendor
- 3.techzine.eu — red hat ai 3 5 focuses on governance and shared gpus
- 4.techstrong.ai — red hat extends scope and reach of ai platform
- 5.channellife.news — red hat ai 3 5 expands ai safety and observability tools
- 6.hn.algolia.com — search
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Genesys Cloud
Agentic orchestration for customer experience — AI agents, employees and systems on one governed platform
VMware Tanzu Platform
Production AI agents inside your own private cloud, with a deny-by-default runtime
SuperApp
Provider-neutral AI workspace putting team chat, documents and every major model in one governed thread
Sirion
AI-native contract lifecycle management where specialist agents draft, redline and govern enterprise contracts