Controlling Reasoning Effort in LLMs
by Sebastian Raschka (Ahead of AI)
How low/medium/high reasoning effort actually gets trained into DeepSeek V4, Kimi, Qwen3, GLM-5, Nemotron and gpt-oss.
Overview
Published July 18, 2026 in Sebastian Raschka's Ahead of AI newsletter, this guide explains how reasoning models are trained and how the 'reasoning effort' setting exposed by model APIs is implemented. It opens with a definition of reasoning models and a short overview of training with reinforcement learning from verifiable rewards (RLVR), where a checker such as SymPy or unit tests rewards only correct final answers, and of the 'aha' moments in which backtracking and self-correction emerge from that reward alone, as in DeepSeek-R1-Zero. It then separates inference scaling (self-consistency, self-refinement) from training scaling. The middle sections cover think tags and format rewards, Qwen3's soft and hard on/off switches (Thinking Mode Fusion and enable_thinking=False), the gpt-oss system-prompt effort line, and how effort level relates to response length and quality. The longest section is a comparison of published recipes: DeepSeek V4 trains Non-think, Think High and Think Max specialists and distills them into one checkpoint; Nemotron 3 Ultra truncates reasoning traces at random budgets during SFT; Kimi K2.5 alternates budgeted and unconstrained RL, and Kimi K3 adds a budget penalty and merges nine specialists through multi-teacher on-policy distillation; GLM-5 adds turn-level, interleaved and preserved thinking; Inkling conditions RL on a continuous 0-to-1 effort value. Raschka is the author of Build a Reasoning Model (From Scratch) and Build a Large Language Model (From Scratch).
At a Glance
- Topic
- Models
- Level
- Intermediate
- Format
- Guide
- Cost
- Free
- Duration
- ~45 min read, 30+ figures
- Provider
- Sebastian Raschka (Ahead of AI)
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓How RLVR trains reasoning models by rewarding only answers a checker verifies
- ✓Why backtracking and self-correction emerge from outcome rewards without rewarding the reasoning trace
- ✓How inference scaling methods like self-consistency and self-refinement differ from training-time scaling
- ✓How Qwen3 implements soft and hard thinking switches, including the empty think block
- ✓How effort levels are trained with length penalties in RL or length-targeted SFT data
- ✓How DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5/K3, GLM-5 and Inkling each implement effort control
- ✓Why changing effort on a fixed model is inference scaling while choosing a bigger model is not
Highlights
- •Side-by-side comparison of effort-control recipes from six open-weight model families, taken from their technical reports
- •Over 30 explanatory figures; the full article was readable without a paid subscription when checked on 2026-10-10
- •Reached 84 points on Hacker News, where commenters called it very high quality and info-dense
- •Links each recipe back to the model's technical report so you can go to the primary source
Who It’s For
Best For
- ✓AI engineers choosing reasoning-effort settings for cost and latency in production
- ✓ML engineers planning post-training runs that need controllable reasoning length
- ✓Developers serving open-weight reasoning models who want to understand thinking toggles and budgets
Prerequisites
- •Working knowledge of how LLMs generate tokens and what SFT and RL post-training are
- •Familiarity with reasoning models such as DeepSeek-R1 or o1 helps but is explained briefly
FAQ
What is Controlling Reasoning Effort in LLMs?
A long illustrated guide from Sebastian Raschka for engineers who pick reasoning-effort settings in production and want to know what those knobs do inside the model. It covers think tokens, on/off reasoning switches and how labs train effort levels, then compares the published recipes of six open-weight model families side by side.
Is Controlling Reasoning Effort in LLMs free?
Controlling Reasoning Effort in LLMs is free to access.
What level is Controlling Reasoning Effort in LLMs for?
Controlling Reasoning Effort in LLMs is aimed at a intermediate audience. Recommended background: Working knowledge of how LLMs generate tokens and what SFT and RL post-training are, Familiarity with reasoning models such as DeepSeek-R1 or o1 helps but is explained briefly.
How long does Controlling Reasoning Effort in LLMs take?
Expect roughly ~45 min read, 30+ figures. Most learners work through it at their own pace.
What will I learn from Controlling Reasoning Effort in LLMs?
You'll learn: How RLVR trains reasoning models by rewarding only answers a checker verifies; Why backtracking and self-correction emerge from outcome rewards without rewarding the reasoning trace; How inference scaling methods like self-consistency and self-refinement differ from training-time scaling; How Qwen3 implements soft and hard thinking switches, including the empty think block; How effort levels are trained with length penalties in RL or length-targeted SFT data; How DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5/K3, GLM-5 and Inkling each implement effort control; Why changing effort on a fixed model is inference scaling while choosing a bigger model is not.
Topics
Sources
This page was written from 3 sources, 1 on domains other than magazine.sebastianraschka.com.