Fine-TuningAgenticModels

OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL

by OpenRLHF

AdvancedDocumentationFreeSelf-paced reference; ~3-4 hours to read the core guides, plus GPU time for the 4-GPU quick-start run

The official guide to training LLMs with PPO, GRPO and REINFORCE++ on Ray and vLLM, from SFT through multi-turn agent RL.

Start LearningAdded Oct 7, 2026 · Updated Oct 7, 2026

Overview

OpenRLHF's documentation is organised into four parts. Getting Started is a Quick Start covering installation (an NVIDIA PyTorch 25.11 Docker image, pip with vLLM 0.19.0 or later, or an editable source install), CLI notes, datasets, pretrained models and the typical workflow. Core Concepts explains the two design ideas: a Ray plus vLLM distributed architecture, where vLLM handles generation and DeepSpeed ZeRO-3 handles training straight from Hugging Face checkpoints, and an agent-based, token-in-token-out execution paradigm that separates single-turn or multi-turn execution from the choice of RL algorithm. Training Guides cover the RL training guide (execution modes, algorithms, optimizer options including Muon, tuning, vision-language RLHF and logging), supervised and preference training (SFT, reward models, DPO) and the full list of CLI options. Scaling and Operations covers the Hybrid Engine, which collocates actor, critic, reward, reference and vLLM on the same GPUs, plus async training with partial rollout, performance tuning, checkpointing, RingAttention sequence parallelism, multi-node training, NVIDIA Docker and troubleshooting. The docs track version 0.10.2, which moved to hierarchical dotted flags such as --actor.model_name_or_path. The canonical first run trains Qwen3-4B-Thinking on math reasoning with REINFORCE++-baseline and a Python reward function on 4 GPUs. The project describes itself in an arXiv paper (Hu et al., 2405.11143) and lists Google, ByteDance, Tencent, Alibaba, Allen AI and the Berkeley Starling team among its users.

At a Glance

Topic
Fine-Tuning
Level
Advanced
Format
Documentation
Cost
Free
Duration
Self-paced reference; ~3-4 hours to read the core guides, plus GPU time for the 4-GPU quick-start run
Provider
OpenRLHF
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • ✓Install OpenRLHF through Docker, pip with vLLM, or an editable source checkout
  • ✓Run the three-stage pipeline of SFT, reward model training and PPO or GRPO
  • ✓Skip the SFT and reward model stages for reasoning tasks using custom Python reward functions
  • ✓Switch between PPO, REINFORCE++, REINFORCE++-baseline, GRPO, Dr. GRPO and RLOO with a single configuration flag
  • ✓Configure single-turn and multi-turn agent execution modes on the token-in-token-out pipeline
  • ✓Use the Hybrid Engine to share GPUs between actor, critic, reward, reference and vLLM
  • ✓Enable async training with partial rollout so generation overlaps with training
  • ✓Scale to multi-node training with checkpointing, RingAttention sequence parallelism and performance tuning

Highlights

  • •Around 10.1k GitHub stars and 1.0k forks under an Apache-2.0 license, with users listed from Google, ByteDance, Tencent, Alibaba and Allen AI
  • •Actively maintained in 2026: ProRL V2 integration in February, VLM and multi-turn VLM RL in April, FlashREINFORCE support in September
  • •One flag switches the RL algorithm, so comparing GRPO against REINFORCE++ or RLOO on the same setup takes little extra work
  • •Covers the production side most RL tutorials skip: resumable checkpoints, off-policy correction, remote HTTP reward models and multi-node runs

Who It’s For

Best For

  • ✓ML engineers moving from supervised fine-tuning to RLHF or RL with verifiable rewards
  • ✓Teams training reasoning or tool-using models on their own multi-GPU clusters
  • ✓Researchers comparing PPO, GRPO and REINFORCE++ variants on the same infrastructure
  • ✓Engineers who need vision-language model RLHF without writing a custom trainer

Prerequisites

  • •Solid PyTorch and Hugging Face Transformers experience, including supervised fine-tuning
  • •Working knowledge of RL for LLMs (policy gradients, PPO or GRPO, reward models)
  • •Access to multiple NVIDIA GPUs (the quick start uses 4; the repo cites 8x A100 80GB configurations)
  • •Comfort with Ray, DeepSpeed and vLLM, or willingness to debug distributed jobs

FAQ

What is OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL?

The official documentation for OpenRLHF, an open-source Apache-2.0 framework for reinforcement learning from human feedback and RL with verifiable rewards. It is written for ML engineers who already fine-tune models and want to run SFT, reward modeling, DPO and RL (PPO, GRPO, REINFORCE++) on their own GPUs, including multi-turn agent rollouts.

Is OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL free?

OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL is free to access.

What level is OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL for?

OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL is aimed at a advanced audience. Recommended background: Solid PyTorch and Hugging Face Transformers experience, including supervised fine-tuning, Working knowledge of RL for LLMs (policy gradients, PPO or GRPO, reward models), Access to multiple NVIDIA GPUs (the quick start uses 4; the repo cites 8x A100 80GB configurations), Comfort with Ray, DeepSpeed and vLLM, or willingness to debug distributed jobs.

How long does OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL take?

Expect roughly Self-paced reference; ~3-4 hours to read the core guides, plus GPU time for the 4-GPU quick-start run. Most learners work through it at their own pace.

What will I learn from OpenRLHF Documentation — Ray + vLLM Framework for RLHF, GRPO and Agent RL?

You'll learn: Install OpenRLHF through Docker, pip with vLLM, or an editable source checkout; Run the three-stage pipeline of SFT, reward model training and PPO or GRPO; Skip the SFT and reward model stages for reasoning tasks using custom Python reward functions; Switch between PPO, REINFORCE++, REINFORCE++-baseline, GRPO, Dr. GRPO and RLOO with a single configuration flag; Configure single-turn and multi-turn agent execution modes on the token-in-token-out pipeline; Use the Hybrid Engine to share GPUs between actor, critic, reward, reference and vLLM; Enable async training with partial rollout so generation overlaps with training; Scale to multi-node training with checkpointing, RingAttention sequence parallelism and performance tuning.

Topics

rlhfgrpopporeinforcement-learningvllmray

Sources

This page was written from 3 sources, 1 on domains other than openrlhf.readthedocs.io.

  1. 1.openrlhf.readthedocs.io — latestvendor
  2. 2.openrlhf.readthedocs.io — quick startvendor
  3. 3.github.com — OpenRLHF