R

Red Hat AI

by Red Hat (IBM)

Enterprise PlatformInfrastructure & CloudGovernance & SecurityAI Agents & Orchestration

Run any model on any accelerator across hybrid cloud, with safety evidence and cost attribution built in

Subscription · Contact for pricing·Added Sep 10, 2026·Updated Sep 10, 2026
Share:
THE DAILY BRIEF
Red Hat AI

by Red Hat (IBM)

Enterprise PlatformInfrastructure & CloudGovernance & SecurityAI Agents & Orchestration

Run any model on any accelerator across hybrid cloud, with safety evidence and cost attribution built in

Subscription · Contact for pricing

Red Hat AI is a hybrid-cloud platform for building, serving and governing enterprise AI workloads on Kubernetes. It bundles Red Hat AI Enterprise, OpenShift AI, RHEL AI and Red Hat AI Inference so platform teams can run any model on any accelerator on-premises, at the edge or across clouds, with pre-deployment safety benchmarking, multi-tenant GPU controls and per-user token metering.

At a Glance

Category
Enterprise Platform
Pricing
Subscription, Contact for pricing
Target Market
CIOs, CTOs, Platform Engineers, MLOps Teams, Enterprise Architects
Deployment
Hybrid, Self-hosted, Multi-cloud, Cloud-first
Founded
1993
Headquarters
Raleigh, United States
Team Size
500+

Key Features

  • EvalHub safety benchmarking
  • Validated model catalog
  • llm-d distributed inference
  • Multi-tenant GPU controls
  • Per-user token metering and observability
  • AutoRAG and agent tooling
  • Hosted control planes on OpenShift Virtualization

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • GPU-as-a-service for internal teams
  • Regulated model approval
  • Sovereign or on-premises inference
  • Multi-cloud inference standardisation
  • Governed internal agents

Ideal For

Best For

  • Enterprises already running OpenShift that want AI serving on the same control plane as their existing workloads
  • Regulated organisations needing auditable pre-deployment safety and compliance evidence for custom models, RAG systems and agents
  • Platform teams chargebacking shared GPU capacity across multiple internal tenants with fair-share scheduling and per-user token metering
  • Hybrid and sovereign deployments where inference must stay on-premises, at the edge or in a specific jurisdiction rather than in a vendor's cloud
  • Teams standardising on open inference components — vLLM and llm-d — to avoid lock-in to a single cloud's model-serving stack

Not Ideal For

  • Small teams or startups that just want an API key: this is heavyweight Kubernetes platform software, and G2 reviewers of the predecessor OpenShift Data Science product consistently name setup and configuration complexity as the main drawback
  • Organisations with no Kubernetes or OpenShift skills in-house — OpenShift-specific concepts like Routes and Operators add a learning curve on top of AI tooling that is itself unfamiliar
  • Buyers who need transparent list pricing: Red Hat publishes subscription guides but not prices, and accelerator subscriptions are quoted per deployment
  • Anyone wanting a turnkey frontier model service; Red Hat serves open and third-party models but does not train a competitive frontier model of its own

Market Analysis

Enterprise-gradeHybrid cloudOpen source

Pros

  • Genuinely hybrid — the same platform and tooling run on-premises, at the edge and on AWS, Azure and CoreWeave, which is rare among enterprise AI platforms and decisive for regulated or sovereign workloads
  • Governance is built in rather than bought separately: pre-deployment safety benchmarking, model risk scoring, guardrails on tool calls and agent tracing all ship in the platform
  • Cost attribution is solved at the platform layer with per-user token metering and showback dashboards, which is what makes internal GPU-as-a-service chargeback workable
  • Backed by Red Hat's subscription support model and IBM ownership, so a platform bet here does not carry the vendor-survival risk of a venture-stage alternative
  • Open inference foundation (vLLM, llm-d) reduces lock-in relative to a hyperscaler's proprietary serving stack

Cons

  • Operational complexity is the recurring theme in independent reviews: G2 reviewers of the predecessor Red Hat OpenShift Data Science product repeatedly cite a steep initial learning curve and difficult setup and configuration, and the broader OpenShift line is described as heavy compared with plain Kubernetes
  • It presumes an OpenShift investment — OpenShift-specific concepts such as Routes and Operators are extra surface area on top of the AI tooling, which makes this a poor fit for smaller teams or simple workloads
  • No published pricing at all, and the self-managed accelerator subscription meters per physical GPU, so cost grows with hardware rather than usage and is hard to model before a quote
  • Practitioner mindshare is thin relative to the hyperscalers: Hacker News threads about OpenShift AI are almost entirely low-engagement press-release reposts, and the most-discussed item was a 2025 root-access vulnerability allowing full cluster takeover rather than a capability
  • Several 3.5 headline items are not production-ready — Amazon EKS support for llm-d is technology preview, vLLM Omni multimodal serving is early access, and storage offloading and the Kubeflow Spark Operator are developer preview

Pricing

Red Hat AI Enterprise

Contact for pricing

  • Unified model and application lifecycle
  • Unified entitlement covering unlimited hardware accelerators per node
  • EvalHub safety benchmarking
  • Validated model catalog

Red Hat OpenShift AI (self-managed)

Contact for pricing

  • Per-AI-Accelerator subscription for each physical GPU, TPU, NPU, FPGA or DPU
  • Runs on customer-managed OpenShift
  • Hybrid cloud, on-premises and edge deployment

Red Hat Enterprise Linux AI

Contact for pricing

  • Single-server LLM deployment
  • Bootable RHEL image with inference runtime

Red Hat publishes subscription guides but no list prices, so every deal is quoted. The metering unit is what matters: self-managed OpenShift AI requires a separate Red Hat AI Accelerator subscription for each discrete physical accelerator card — GPU, TPU, NPU, FPGA or DPU — that is not part of the CPU package, at the same rate regardless of OpenShift edition, so cost scales with GPU count rather than users or tokens. Red Hat AI Enterprise instead provides a unified entitlement covering unlimited accelerators per node, which changes the arithmetic sharply for dense GPU servers. Trials are offered; there is no free tier.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Red Hat AI is a hybrid-cloud platform for building, serving and governing enterprise AI workloads on Kubernetes. It bundles Red Hat AI Enterprise, OpenShift AI, RHEL AI and Red Hat AI Inference so platform teams can run any model on any accelerator on-premises, at the edge or across clouds, with pre-deployment safety benchmarking, multi-tenant GPU controls and per-user token metering.

Red Hat AI is Red Hat's umbrella enterprise AI platform, built on OpenShift and positioned around a single claim: any model, any agent, any hardware accelerator, anywhere. It comprises Red Hat AI Enterprise (a unified model and application lifecycle environment), Red Hat OpenShift AI (predictive and generative model lifecycle at scale on Kubernetes), Red Hat Enterprise Linux AI for single-server LLM deployment, and Red Hat AI Inference, which is powered by vLLM. Version 3.5, generally available 9 September 2026, moved the platform's centre of gravity from getting models into production to running them as governed infrastructure. Its headline addition is EvalHub, now GA: an evaluation orchestration service with a versioned REST API that runs multi-framework benchmarks inside isolated, horizontally scalable Kubernetes jobs, letting teams pre-screen custom models, RAG configurations and agents for risk and emit auditable compliance reports. The validated model catalog gained 20-plus entries including Google Gemma 4, NVIDIA Nemotron 3 and Alibaba Qwen, each carrying Garak security scores and PII-exposure and toxicity assessments. On infrastructure, 3.5 added fair-share GPU scheduling, priority-aware serving that protects real-time inference from background jobs, dynamic GPU capacity borrowing, and hosted control planes on OpenShift Virtualization for VM-level tenant isolation on shared hardware. The llm-d distributed inference scheduler went generally available on CoreWeave CKS and Microsoft Azure with Amazon EKS in technology preview — the first time it runs beyond OpenShift. For agents, the Responses API and AutoRAG reached GA with pgvector and multilingual document support, NVIDIA NeMo Guardrails intercept malicious tool calls, and AI Hub ships starter templates for code review, document processing and research. Observability added GPU utilisation and inference-health dashboards, MLflow visual agentic tracing, and per-user token metering for model-as-a-service showback.

Ideal Buyer

The platform engineering or MLOps leader at an enterprise already standardised on OpenShift or Kubernetes, who must serve multiple internal AI tenants from shared GPUs and produce audit evidence for what those models and agents do.

Key Benefit

One governed control plane for models, agents and GPUs that works identically on-premises, at the edge and on AWS, Azure or CoreWeave — with pre-deployment safety scores and per-user token showback as first-class outputs rather than bolt-ons.

At a Glance

Category
Enterprise Platform
Pricing
Subscription, Contact for pricing
Target Market
CIOs, CTOs, Platform Engineers, MLOps Teams, Enterprise Architects
Deployment
Hybrid, Self-hosted, Multi-cloud, Cloud-first
Founded
1993
Headquarters
Raleigh, United States
Team Size
500+

Key Features

  • EvalHub safety benchmarking

    Generally available evaluation orchestration service that runs multi-framework benchmarks in isolated Kubernetes jobs and emits auditable compliance reports before deployment.

  • Validated model catalog

    Over twenty vetted models including Gemma 4, Nemotron 3 and Qwen, each shipping Garak security scores plus PII-exposure and toxicity risk assessments.

  • llm-d distributed inference

    Distributed inference scheduler now generally available on CoreWeave CKS and Microsoft Azure, with Amazon EKS in technology preview and new KV-cache reuse scorers.

  • Multi-tenant GPU controls

    Fair-share scheduling, priority-aware serving and dynamic capacity borrowing let several tenants share GPUs without background jobs starving real-time inference.

  • Per-user token metering and observability

    Dashboards for GPU utilisation, inference health and model performance, plus token metering that makes model-as-a-service chargeback and showback possible.

  • AutoRAG and agent tooling

    GA Responses API, pgvector-backed multilingual RAG, NeMo Guardrails that intercept malicious tool calls, and starter templates for code review and document workflows.

  • Hosted control planes on OpenShift Virtualization

    Gives each tenant a dedicated cluster control plane with VM-level isolation while still sharing the underlying accelerator hardware.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • GPU-as-a-service for internal teams

    A platform team sells shared accelerator capacity internally, using fair-share scheduling to prevent contention and token metering to attribute cost per department.

  • Regulated model approval

    A risk function runs EvalHub against a fine-tuned model and its RAG pipeline, producing safety and PII-exposure evidence before the model reaches production.

  • Sovereign or on-premises inference

    An organisation that cannot send prompts to a public API serves open models in its own datacentre while keeping the same tooling as its cloud footprint.

  • Multi-cloud inference standardisation

    A team runs llm-d on CoreWeave and Azure alongside OpenShift on-premises, keeping one scheduler and one operational model across all three environments.

  • Governed internal agents

    Developers build document-processing or code-review agents from AI Hub templates while NeMo Guardrails block malicious tool calls and MLflow traces every agent step.

Ideal For

Best For

  • Enterprises already running OpenShift that want AI serving on the same control plane as their existing workloads
  • Regulated organisations needing auditable pre-deployment safety and compliance evidence for custom models, RAG systems and agents
  • Platform teams chargebacking shared GPU capacity across multiple internal tenants with fair-share scheduling and per-user token metering
  • Hybrid and sovereign deployments where inference must stay on-premises, at the edge or in a specific jurisdiction rather than in a vendor's cloud
  • Teams standardising on open inference components — vLLM and llm-d — to avoid lock-in to a single cloud's model-serving stack

Not Ideal For

  • Small teams or startups that just want an API key: this is heavyweight Kubernetes platform software, and G2 reviewers of the predecessor OpenShift Data Science product consistently name setup and configuration complexity as the main drawback
  • Organisations with no Kubernetes or OpenShift skills in-house — OpenShift-specific concepts like Routes and Operators add a learning curve on top of AI tooling that is itself unfamiliar
  • Buyers who need transparent list pricing: Red Hat publishes subscription guides but not prices, and accelerator subscriptions are quoted per deployment
  • Anyone wanting a turnkey frontier model service; Red Hat serves open and third-party models but does not train a competitive frontier model of its own

Integrations

SDK Available
SDK:Python

Deployment

On-Premise

Market Analysis

Enterprise-gradeHybrid cloudOpen source

Pros

  • Genuinely hybrid — the same platform and tooling run on-premises, at the edge and on AWS, Azure and CoreWeave, which is rare among enterprise AI platforms and decisive for regulated or sovereign workloads
  • Governance is built in rather than bought separately: pre-deployment safety benchmarking, model risk scoring, guardrails on tool calls and agent tracing all ship in the platform
  • Cost attribution is solved at the platform layer with per-user token metering and showback dashboards, which is what makes internal GPU-as-a-service chargeback workable
  • Backed by Red Hat's subscription support model and IBM ownership, so a platform bet here does not carry the vendor-survival risk of a venture-stage alternative
  • Open inference foundation (vLLM, llm-d) reduces lock-in relative to a hyperscaler's proprietary serving stack

Cons

  • Operational complexity is the recurring theme in independent reviews: G2 reviewers of the predecessor Red Hat OpenShift Data Science product repeatedly cite a steep initial learning curve and difficult setup and configuration, and the broader OpenShift line is described as heavy compared with plain Kubernetes
  • It presumes an OpenShift investment — OpenShift-specific concepts such as Routes and Operators are extra surface area on top of the AI tooling, which makes this a poor fit for smaller teams or simple workloads
  • No published pricing at all, and the self-managed accelerator subscription meters per physical GPU, so cost grows with hardware rather than usage and is hard to model before a quote
  • Practitioner mindshare is thin relative to the hyperscalers: Hacker News threads about OpenShift AI are almost entirely low-engagement press-release reposts, and the most-discussed item was a 2025 root-access vulnerability allowing full cluster takeover rather than a capability
  • Several 3.5 headline items are not production-ready — Amazon EKS support for llm-d is technology preview, vLLM Omni multimodal serving is early access, and storage offloading and the Kubeflow Spark Operator are developer preview

Pricing

Free Trial Available

Red Hat AI Enterprise

Contact for pricing

  • Unified model and application lifecycle
  • Unified entitlement covering unlimited hardware accelerators per node
  • EvalHub safety benchmarking
  • Validated model catalog

Red Hat OpenShift AI (self-managed)

Contact for pricing

  • Per-AI-Accelerator subscription for each physical GPU, TPU, NPU, FPGA or DPU
  • Runs on customer-managed OpenShift
  • Hybrid cloud, on-premises and edge deployment

Red Hat Enterprise Linux AI

Contact for pricing

  • Single-server LLM deployment
  • Bootable RHEL image with inference runtime

Red Hat publishes subscription guides but no list prices, so every deal is quoted. The metering unit is what matters: self-managed OpenShift AI requires a separate Red Hat AI Accelerator subscription for each discrete physical accelerator card — GPU, TPU, NPU, FPGA or DPU — that is not part of the CPU package, at the same rate regardless of OpenShift edition, so cost scales with GPU count rather than users or tokens. Red Hat AI Enterprise instead provides a unified entitlement covering unlimited accelerators per node, which changes the arithmetic sharply for dense GPU servers. Trials are offered; there is no free tier.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Sources

This page was written from 6 sources, 4 on domains other than redhat.com.

  1. 1.redhat.comaivendor
  2. 2.redhat.comred hat puts safety and observability core enterprise ai redvendor
  3. 3.techzine.eured hat ai 3 5 focuses on governance and shared gpus
  4. 4.techstrong.aired hat extends scope and reach of ai platform
  5. 5.channellife.newsred hat ai 3 5 expands ai safety and observability tools
  6. 6.hn.algolia.comsearch
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe