Fine-TuningModelsML

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

by Sina Weibo (WeiboAI)

AdvancedPaperFree14-page report, ~1-2 hours to read

The post-training recipe that took Qwen2.5-Coder-3B to 94.3 on AIME26: curriculum SFT, MGPO reinforcement learning and offline self-distillation.

Start LearningAdded Oct 9, 2026 · Updated Oct 9, 2026

Overview

VibeThinker-3B is a 14-page technical report by Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou and Junlin Zhang of Sina Weibo, posted to arXiv on June 15, 2026. It extends their November 2025 VibeThinker-1.5B work. The model is post-trained from Qwen2.5-Coder-3B with what the authors call the Spectrum-to-Signal approach, and the Methods section lays it out in order. Section 2.1 covers supervised fine-tuning: seed queries limited to those with reliable answers or unit tests, several teacher traces per query with majority-vote pseudo-labels, n-gram and LLM filtering, then a two-stage curriculum (5 epochs on the full set, 2 epochs on a hard subset of traces over 5K tokens) and Diversity-Exploring Distillation, which merges domain specialist checkpoints chosen by Pass@K. Section 2.2 covers reinforcement learning with MaxEnt-Guided Policy Optimization (MGPO), a GRPO-style objective that up-weights prompts with roughly 50% group accuracy, run in sequential Math, Code and STEM stages at 64K context, plus a Long2Short stage that rewards shorter correct answers. Sections 2.3 and 2.4 cover offline self-distillation into one student and instruction-following RL. Reported scores are 94.3 on AIME26 (97.1 with test-time scaling), 80.2 Pass@1 on LiveCodeBench v6 and 93.4 on IFEval; GPQA-Diamond is 70.2. Weights are MIT-licensed on Hugging Face.

At a Glance

Topic
Fine-Tuning
Level
Advanced
Format
Paper
Cost
Free
Duration
14-page report, ~1-2 hours to read
Provider
Sina Weibo (WeiboAI)
Hands-on
No
Certificate
None

What You’ll Learn

  • ✓How to build verified SFT data with teacher traces, majority-vote labels and sandbox checks
  • ✓How a two-stage SFT curriculum moves from the full dataset to long, hard reasoning traces
  • ✓How Diversity-Exploring Distillation merges domain specialist checkpoints selected by Pass@K
  • ✓How MGPO modifies a GRPO-style objective to focus RL on prompts near 50% group accuracy
  • ✓How to sequence math, code and STEM RL stages with verifiable rewards at 64K context
  • ✓How a zero-sum length reward in a Long2Short stage shortens correct reasoning traces
  • ✓How offline self-distillation folds several domain RL checkpoints back into one student model

Highlights

  • •A complete small-model reasoning recipe with concrete hyperparameters (learning rates, epochs, batch size, the λ = 0.2 length shift)
  • •Released weights under the MIT license; community GGUF quantizations appeared within a day and Hacker News users ran it on an RTX 2070 Super at about 110 tokens/s
  • •The 398-point, 205-comment Hacker News thread is a useful reality check: strong on math and self-contained Python, weak at tool calling, structured output and security review
  • •The authors guard against contamination with April to May 2026 LeetCode contests, passing 123 of 128 first-attempt submissions, and limit their claims to verifiable reasoning

Who It’s For

Best For

  • ✓ML engineers post-training small open models for math or coding reasoning
  • ✓Teams weighing a cheap specialist reasoning model under a larger orchestrator
  • ✓Researchers comparing GRPO variants and curriculum SFT designs
  • ✓Practitioners who want a reproducible recipe rather than a leaderboard number

Prerequisites

  • •Solid grounding in supervised fine-tuning and policy-gradient RL (PPO or GRPO)
  • •Familiarity with reasoning benchmarks such as AIME and LiveCodeBench
  • •Comfort reading ML papers with training hyperparameters and ablations

FAQ

What is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models?

The VibeThinker-3B technical report is a June 2026 paper from Sina Weibo for engineers who post-train small open models for reasoning. It documents a full pipeline on a 3B base, from SFT data construction and a two-stage curriculum through multi-domain RL, a length-shortening stage and self-distillation, so you can adapt the recipe to your own small reasoning model.

Is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models free?

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models is free to access.

What level is VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models for?

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models is aimed at a advanced audience. Recommended background: Solid grounding in supervised fine-tuning and policy-gradient RL (PPO or GRPO), Familiarity with reasoning benchmarks such as AIME and LiveCodeBench, Comfort reading ML papers with training hyperparameters and ablations.

How long does VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models take?

Expect roughly 14-page report, ~1-2 hours to read. Most learners work through it at their own pace.

What will I learn from VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models?

You'll learn: How to build verified SFT data with teacher traces, majority-vote labels and sandbox checks; How a two-stage SFT curriculum moves from the full dataset to long, hard reasoning traces; How Diversity-Exploring Distillation merges domain specialist checkpoints selected by Pass@K; How MGPO modifies a GRPO-style objective to focus RL on prompts near 50% group accuracy; How to sequence math, code and STEM RL stages with verifiable rewards at 64K context; How a zero-sum length reward in a Long2Short stage shortens correct reasoning traces; How offline self-distillation folds several domain RL checkpoints back into one student model.

Topics

reasoning modelsgrpopost-trainingsmall language modelsdistillationreinforcement learning

Sources

This page was written from 4 sources, 2 on domains other than arxiv.org.

  1. 1.arxiv.org — 2606.16140vendor
  2. 2.arxiv.org — 2606.16140v1vendor
  3. 3.hn.algolia.com — 48639240
  4. 4.venturebeat.com — why weibos tiny vibethinker 3b has the ai world arguing over