Skip to content

Repository files navigation

Verifier-Induced Support Reshaping in On-Policy Optimization

Verifier-Induced Support Reshaping in On-Policy Optimization

Shaohang Wei1‡, Zikun Su2, Feifan Song1, Wen Luo1, Wei Li1, Guangyue Peng1, Houfeng Wang1†

1Peking University   2BUPT
‡ Project Lead   † Corresponding Author
Correspondence: wanghf@pku.edu.cn

Paper on arXiv Project website Code license: Apache 2.0 Python 3.10 or newer

Project Website ↗  ·  Overview  ·  Results  ·  Findings  ·  Getting started  ·  Citation

Overview

Can a model still discover successful behaviors for its next training objective? VISR studies how on-policy reinforcement learning with verifiable rewards (RLVR) changes this ability across mathematical reasoning and instruction following. We define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Improving the current objective can make those trajectories harder to sample and reinforce in a later stage.

Overview: Math-RLVR and IF-RLVR reshape the successful trajectories available to later on-policy training.

Read the paper · Explore the project · Static SVG · Figure PDF

Main results

Math-RLVR improves average instruction-following success while reducing coverage under repeated sampling. The pattern appears in both model families on IFEval and IFBench. Here, pass@1 measures average rollout success; best@32 measures the fraction of prompts with at least one successful response among 32 samples.

Changes from Base to the final Math-RLVR checkpoint, on the percentage scale:

Model Benchmark Δ pass@1 Δ best@32
Qwen3-8B-Base IFEval +6.5% −9.8%
Qwen2.5-Math-7B IFEval +7.9% −11.4%
Qwen3-8B-Base IFBench +3.2% −6.7%
Qwen2.5-Math-7B IFBench +1.6% −3.7%

In the reverse direction, IF-RLVR lowers math searchability. On AIME, best@k decreases at every tested budget (k = 4, 8, 16, 32), while visible response openings shift from step-by-step reasoning toward direct answers.

View the training curves and opening-route analysis

Math-RLVR training curves: average IF success rises while best@32 falls.

IF-RLVR training curves: math best@k decreases as visible opening routes change.

Explore the full results

Response openings and math searchability

Distribution shifts are largest at the first response token. Controlled opening interventions improve math best@32 from IF-RLVR checkpoints.

The first response token has the largest mean distribution shift in every tested model, verifier, and benchmark combination. Forcing Base-side or deliberative openings from IF-RLVR checkpoints improves best@32 on AIME and MATH-500 in both model families. These controlled interventions show that response openings affect math searchability in the tested settings.

Position sweeps and intervention details

Key findings

  • Average success and sampling coverage can move in opposite directions. A higher pass@1 can coexist with fewer prompts yielding any successful response within a fixed budget.
  • Response openings matter for later search. Token-distribution measurements and controlled interventions identify the opening as a point where RLVR changes which successful responses remain reachable.
  • Preservation remains partial. Reference-policy constraints and opening priors provide limited retention in the tested settings; on-policy distillation outcomes depend on the teacher checkpoint.

Getting started

The repository includes the training framework, evaluation pipeline, analysis scripts, and selected paper tables. Paper checkpoints and raw rollouts are not bundled; see resource availability.

git clone https://github.com/sylvain-wei/VISR.git
cd VISR
Task Guide
Install dependencies and configure models / data Evaluation setup
Run a small check or the full evaluation Evaluation commands
Inspect training recipes and their requirements Training scope · DAPO reference recipe
Inspect analyses and released tables Analysis guide · Paper tables

After installing the evaluation environment and configuring checkpoint and dataset paths:

cd eval
python scripts/check_env.py --strict
bash scripts/dry_run.sh
bash scripts/run_rq1_required.sh

License and acknowledgments

Code and repository documentation use the Apache License 2.0; paper figures retain their CC BY 4.0 license. We build on verl and evaluation components from Google Research IFEval and AllenAI IFBench. See third-party notices and asset credits for attribution and license details.

Citation

@misc{wei2026verifier,
  title         = {Verifier-Induced Support Reshaping in On-Policy Optimization},
  author        = {Shaohang Wei and Zikun Su and Feifan Song and Wen Luo and Wei Li and Guangyue Peng and Houfeng Wang},
  year          = {2026},
  eprint        = {2608.00220},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2608.00220},
  url           = {https://arxiv.org/abs/2608.00220}
}

About

[UnderReview'27] Verifier-Induced Support Reshaping in On-Policy Optimization

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages