VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
by Sina Weibo (WeiboAI)
The post-training recipe that took Qwen2.5-Coder-3B to 94.3 on AIME26: curriculum SFT, MGPO reinforcement learning and offline self-distillation.
Overview
VibeThinker-3B is a 14-page technical report by Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou and Junlin Zhang of Sina Weibo, posted to arXiv on June 15, 2026. It extends their November 2025 VibeThinker-1.5B work. The model is post-trained from Qwen2.5-Coder-3B with what the authors call the Spectrum-to-Signal approach, and the Methods section lays it out in order. Section 2.1 covers supervised fine-tuning: seed queries limited to those with reliable answers or unit tests, several teacher traces per query with majority-vote pseudo-labels, n-gram and LLM filtering, then a two-stage curriculum (5 epochs on the full set, 2 epochs on a hard subset of traces over 5K tokens) and Diversity-Exploring Distillation, which merges domain specialist checkpoints chosen by Pass@K. Section 2.2 covers reinforcement learning with MaxEnt-Guided Policy Optimization (MGPO), a GRPO-style objective that up-weights prompts with roughly 50% group accuracy, run in sequential Math, Code and STEM stages at 64K context, plus a Long2Short stage that rewards shorter correct answers. Sections 2.3 and 2.4 cover offline self-distillation into one student and instruction-following RL. Reported scores are 94.3 on AIME26 (97.1 with test-time scaling), 80.2 Pass@1 on LiveCodeBench v6 and 93.4 on IFEval; GPQA-Diamond is 70.2. Weights are MIT-licensed on Hugging Face.
At a Glance
- Topic
- Fine-Tuning
- Level
- Advanced
- Format
- Paper
- Cost
- Free
- Duration
- 14-page report, ~1-2 hours to read
- Provider
- Sina Weibo (WeiboAI)
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓How to build verified SFT data with teacher traces, majority-vote labels and sandbox checks
- ✓How a two-stage SFT curriculum moves from the full dataset to long, hard reasoning traces
- ✓How Diversity-Exploring Distillation merges domain specialist checkpoints selected by Pass@K
- ✓How MGPO modifies a GRPO-style objective to focus RL on prompts near 50% group accuracy
- ✓How to sequence math, code and STEM RL stages with verifiable rewards at 64K context
- ✓How a zero-sum length reward in a Long2Short stage shortens correct reasoning traces
- ✓How offline self-distillation folds several domain RL checkpoints back into one student model
Highlights
- •A complete small-model reasoning recipe with concrete hyperparameters (learning rates, epochs, batch size, the λ = 0.2 length shift)
- •Released weights under the MIT license; community GGUF quantizations appeared within a day and Hacker News users ran it on an RTX 2070 Super at about 110 tokens/s
- •The 398-point, 205-comment Hacker News thread is a useful reality check: strong on math and self-contained Python, weak at tool calling, structured output and security review
- •The authors guard against contamination with April to May 2026 LeetCode contests, passing 123 of 128 first-attempt submissions, and limit their claims to verifiable reasoning
Who It’s For
Best For
- ✓ML engineers post-training small open models for math or coding reasoning
- ✓Teams weighing a cheap specialist reasoning model under a larger orchestrator
- ✓Researchers comparing GRPO variants and curriculum SFT designs
- ✓Practitioners who want a reproducible recipe rather than a leaderboard number
Prerequisites
- •Solid grounding in supervised fine-tuning and policy-gradient RL (PPO or GRPO)
- •Familiarity with reasoning benchmarks such as AIME and LiveCodeBench
- •Comfort reading ML papers with training hyperparameters and ablations
FAQ
What is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models?
The VibeThinker-3B technical report is a June 2026 paper from Sina Weibo for engineers who post-train small open models for reasoning. It documents a full pipeline on a 3B base, from SFT data construction and a two-stage curriculum through multi-domain RL, a length-shortening stage and self-distillation, so you can adapt the recipe to your own small reasoning model.
Is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models free?
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models is free to access.
What level is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models for?
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models is aimed at a advanced audience. Recommended background: Solid grounding in supervised fine-tuning and policy-gradient RL (PPO or GRPO), Familiarity with reasoning benchmarks such as AIME and LiveCodeBench, Comfort reading ML papers with training hyperparameters and ablations.
How long does VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models take?
Expect roughly 14-page report, ~1-2 hours to read. Most learners work through it at their own pace.
What will I learn from VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models?
You'll learn: How to build verified SFT data with teacher traces, majority-vote labels and sandbox checks; How a two-stage SFT curriculum moves from the full dataset to long, hard reasoning traces; How Diversity-Exploring Distillation merges domain specialist checkpoints selected by Pass@K; How MGPO modifies a GRPO-style objective to focus RL on prompts near 50% group accuracy; How to sequence math, code and STEM RL stages with verifiable rewards at 64K context; How a zero-sum length reward in a Long2Short stage shortens correct reasoning traces; How offline self-distillation folds several domain RL checkpoints back into one student model.
Topics
Sources
This page was written from 4 sources, 2 on domains other than arxiv.org.