AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents
by AReaL Team (inclusionAI)
Train tool-using agents with fully asynchronous RL by pointing an existing OpenAI Agents SDK or CAMEL agent at AReaL's proxy.
Overview
AReaL's documentation is organised as a Tutorial section (installation for NVIDIA GPUs and Huawei Ascend NPUs, a Quickstart, Agentic Reinforcement Learning, Online RL Training, Evaluation, Fine-tuning Large MoE Models, the PyTorch-native Archon training engine and Configurations), a Code Walkthrough that runs GRPO on the GSM8K dataset, Best Practices (diagnosing RL performance, writing agent workflows, debugging, handling OOM issues and performance profiling), Customization (datasets and custom agent workflows), an Algorithms section (asynchronous RL, on-policy distillation, DPO, PPO and GRPO-family methods, M2PO, proximal log-probability approximation and process rewards) and a Reference section covering checkpointing, metrics, allocation mode, LoRA, the Megatron-HF bridge, tree training and the RolloutWorkflow and Agent Workflow APIs. The agentic RL tutorial is the distinctive part: AReaL runs an HTTP proxy in front of its SGLang or vLLM inference engine, so an agent written with the OpenAI Agents SDK or CAMEL-AI can be trained without changing its code, while the proxy records token-level data, builds conversation trees and propagates rewards back through multi-turn dialogues. It runs on local, Slurm or Ray schedulers. The design comes from the AReaL paper (Fu et al., arXiv 2505.24298, revised March 2026), which reports up to 2.77x training speedup over synchronous systems using decoupled rollout workers and a staleness-enhanced PPO variant. AReaL 2.0, released July 2026, moved to a microservice architecture with separate training, inference, agent and weight-update services.
At a Glance
- Topic
- Agentic
- Level
- Advanced
- Format
- Documentation
- Cost
- Free
- Duration
- Self-paced; ~4-5 hours across the tutorials, code walkthrough and best-practice guides, plus cluster time for training runs
- Provider
- AReaL Team (inclusionAI)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓Install AReaL on NVIDIA GPUs or Huawei Ascend NPUs and run the Quickstart
- ✓Walk through a complete GRPO training run on the GSM8K math dataset
- ✓Train an OpenAI Agents SDK or CAMEL-AI agent through AReaL's HTTP proxy without rewriting it
- ✓Write custom reward functions and propagate rewards through multi-turn conversations with discount factors
- ✓Understand asynchronous RL, where rollout workers keep generating while training proceeds on slightly stale samples
- ✓Choose between PPO, GRPO, DAPO, M2PO, on-policy distillation and DPO for a given task
- ✓Diagnose poor RL performance, profile throughput and fix out-of-memory failures on large runs
- ✓Fine-tune large MoE models using Megatron, PyTorch FSDP or the Archon engine
Highlights
- •The proxy approach trains agents built in existing frameworks (OpenAI Agents SDK, CAMEL-AI, Claude Agent SDK) as they are, which most RL libraries do not attempt
- •Backed by a peer-reviewable paper reporting up to 2.77x speedup over synchronous RL at matched or better final performance
- •About 5.8k GitHub stars and steady 2026 releases: TensorRT-LLM integration in April, KPop and IcePop in June, AReaL 2.0 in July
- •Dedicated best-practice pages on diagnosing RL runs, debugging and OOM handling, the problems that actually stall agent RL projects
Who It’s For
Best For
- ✓Engineers who have a working agent and want to improve it with reinforcement learning
- ✓Teams training math, coding, search or customer-service agents on multi-GPU clusters
- ✓Researchers studying asynchronous RL, staleness control and off-policy corrections
- ✓Organisations running RL on Huawei Ascend hardware as well as NVIDIA GPUs
Prerequisites
- •Strong Python and PyTorch skills plus experience fine-tuning LLMs
- •Familiarity with RL for LLMs (PPO, GRPO, reward design)
- •Experience building agents with a framework such as the OpenAI Agents SDK
- •Multi-GPU or multi-node hardware and comfort with Ray or Slurm schedulers
FAQ
What is AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents?
The official documentation for AReaL, an open-source Apache-2.0 system for asynchronous reinforcement learning on LLMs and agents. It is aimed at engineers who already build agents and want to train them with RL, covering installation, a GRPO walkthrough on GSM8K, agentic and online RL tutorials, debugging guides and the algorithms the system ships.
Is AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents free?
AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents is free to access.
What level is AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents for?
AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents is aimed at a advanced audience. Recommended background: Strong Python and PyTorch skills plus experience fine-tuning LLMs, Familiarity with RL for LLMs (PPO, GRPO, reward design), Experience building agents with a framework such as the OpenAI Agents SDK, Multi-GPU or multi-node hardware and comfort with Ray or Slurm schedulers.
How long does AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents take?
Expect roughly Self-paced; ~4-5 hours across the tutorials, code walkthrough and best-practice guides, plus cluster time for training runs. Most learners work through it at their own pace.
What will I learn from AReaL Documentation — Asynchronous Reinforcement Learning for LLM Agents?
You'll learn: Install AReaL on NVIDIA GPUs or Huawei Ascend NPUs and run the Quickstart; Walk through a complete GRPO training run on the GSM8K math dataset; Train an OpenAI Agents SDK or CAMEL-AI agent through AReaL's HTTP proxy without rewriting it; Write custom reward functions and propagate rewards through multi-turn conversations with discount factors; Understand asynchronous RL, where rollout workers keep generating while training proceeds on slightly stale samples; Choose between PPO, GRPO, DAPO, M2PO, on-policy distillation and DPO for a given task; Diagnose poor RL performance, profile throughput and fix out-of-memory failures on large runs; Fine-tune large MoE models using Megatron, PyTorch FSDP or the Archon engine.
Topics
Sources
This page was written from 4 sources, 2 on domains other than areal-ai.io.