---
title: HDPO Hybrid Distillation Policy Optimization
arxiv_id: 2603.23871
source: https://arxiv.org/html/2603.23871
---

 

 HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation 

Report GitHub Issue

 × 

 Title: 

Content selection saved. Describe the issue below:

 Description: 

 Submit without GitHub 
 Submit in GitHub 

 Back to arXiv 

 Why HTML? 

 Report Issue 

 Back to Abstract 

 Download PDF 

 Abstract 

 1 Introduction 

 2 Background 

 2.1 Reinforcement Learning for Language
Models 

 2.2 Group Relative Policy
Optimization 

 2.3 The Cliff Problem 

 2.4 Approaches to the Cliff
Problem 

 2.5 Knowledge Distillation for
Reasoning 

 3 Method: Hybrid Distillation Policy
Optimization 

 3.1 Motivation 

 3.2 Formal Definition 

 Formal Definition. 

 3.3 Algorithm 

 3.4 Theoretical Analysis 

 Proof sketch. 

 Interpretation. 

 Relation to prior work. 

 Proof sketch. 

 Connection to HDPO. 

 4 Experiments 

 4.1 Experimental Setup 

 4.2 Results 

 5 Discussion 

 6 Limitations 

 7 Conclusion 

 A Realizability Gap Bound — Full
Proof 

 A.1 Setup 

 A.2 Lemma 1 (KL Divergence under Logit
Perturbation) 

 A.3 Assumption 1 (Local Lipschitz
Continuity) 

 A.4 Proposition 1 (Realizability Gap
Comparison) 

 A.4.1 Part I – Same-Model
Bound 

 A.4.2 Part II – Cross-Model
Comparison 

 A.5 Remark 

 B Proposition 2 Full
Proof 

 C Experimental
Details 

 D Hardware Variation 

 References 

 License: CC BY 4.0

arXiv:2603.23871v1 [cs.LG] 25 Mar 2026

HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation

 Ken Ding 

NVIDIA 

 kennethd@nvidia.com 

Abstract

Large language models trained with reinforcement learning (RL) for mathematical reasoning face a fundamental challenge: on problems the model cannot solve at all—“cliff” prompts—the RL gradient vanishes entirely, preventing any learning signal from reaching these failure modes. We introduce Hybrid Distillation Policy Optimization (HDPO), which augments standard RL with privileged self-distillation targeting cliff prompts. On each training step, HDPO identifies prompts where all rollouts fail, generates privileged rollouts by providing the model with ground-truth information, filters for correct solutions, and distills the teacher’s token-level distribution into the student. Because teacher and student share the same weights—differing only in their input—the realizability gap is provably bounded, unlike cross-model distillation. We prove that R = 1 R{=}1 filtered privileged generation recovers the optimal KL-regularized RL policy in the hard-threshold limit. Experiments on OpenMathInstruct-2 with Qwen2.5-Math-1.5B-Instruct show that HDPO consistently improves coverage metrics ( pass@4 by + 0.8 +0.8 – 1.1 % 1.1\% , pass@8 by + 0.4 +0.4 – 1.7 % 1.7\% ) while maintaining greedy accuracy, with the distillation weight λ \lambda providing direct control over the exploration–exploitation tradeoff.

 1 Introduction

Recent advances in reinforcement learning from verifiable rewards (RLVR)
have enabled large language models to develop sophisticated mathematical
reasoning capabilities (Shao et al., 2024 ; DeepSeek-AI et al., 2025 ) . Among RL algorithms for language models,
Group Relative Policy Optimization (GRPO)
 (Shao et al., 2024 ) has emerged as a practical and
effective approach, eliminating the need for a critic network by
normalizing rewards within each group of rollouts to compute advantages.

However, GRPO and related policy gradient methods share a fundamental
limitation: they can only learn from problems where at least some
rollouts succeed. When all rollouts for a prompt receive zero reward —
a condition commonly called a "cliff" (Zhang et al., 2025b ) — the advantage estimates are identical
for all trajectories and the policy gradient vanishes. These
zero-gradient prompts are precisely the ones where learning is most
needed, yet they receive no training signal. Le et al.
 (Le et al., 2025 ) characterize this as the zero-variance
prompt problem and show that, on their benchmark, a significant fraction
of prompts falls in this regime—the exact proportion depends on the
dataset and model, but the phenomenon creates a persistent learning dead
zone at the frontier of the model’s capability.

We take a different approach inspired by the principle of learning using
privileged information (Vapnik and Vashist, 2009 ) . In our framework,
Hybrid Distillation Policy Optimization (HDPO), the model acts as both
teacher and student. As the teacher, it receives the problem along with
ground-truth information and generates rollouts from this privileged
context. As the student, it receives only the original problem. Because
both roles use the same model weights, the gap between their
distributions is provably bounded (Proposition 1), unlike cross-model distillation where the gap depends on architectural differences between teacher and student.

HDPO operates as follows: on each training step, after standard GRPO
updates, we identify cliff prompts where all rollouts scored zero. For
these prompts, the model generates privileged rollouts conditioned on
ground truth, filters for correct ones ( R = 1 R\ =\ 1 ), and distills the
teacher’s token-level distribution into the student via JSD. The
 R = 1 R\ =\ 1 filter is not a heuristic — we prove it implements
rejection sampling from the KL-regularized RL optimal policy
(Proposition 2).

Our contributions are as follows:

(1) We introduce HDPO, a hybrid training objective that combines RL with
privileged self-distillation targeting cliff prompts where the RL
gradient vanishes.

(2) We prove that same-model privileged distillation achieves a strictly
tighter realizability gap than cross-model distillation — the gap
depends only on the information content of the ground truth, eliminating
the model-mismatch term that cross-model distillation cannot avoid
(Proposition 1).

(3) We prove that R = 1 R\ =\ 1 filtered privileged generation recovers
the optimal KL-regularized RL policy (Proposition 2), providing
theoretical justification for the teacher construction.

(4) We demonstrate on OpenMathInstruct-2
 (Toshniwal et al., 2024 ) that HDPO improves coverage
( p ​ a ​ s ​ s ​ @ ​ 4 pass@4 , p ​ a ​ s ​ s ​ @ ​ 8 pass@8 ) while maintaining greedy accuracy ( p ​ a ​ s ​ s ​ @ ​ 1 pass@1 ),
with the distillation weight λ \lambda providing explicit control over
the exploration–exploitation tradeoff.

 2 Background

 2.1 Reinforcement Learning for Language
Models

Reinforcement learning-based training of language models formulates text
generation as a sequential decision problem where the policy
 π θ \pi_{\theta} generates tokens autoregressively and receives a scalar
reward upon completion. This framework, popularized by RLHF
 (Ouyang et al., 2022 ) for instruction following, has been
adapted to mathematical reasoning as reinforcement learning from
verifiable rewards (RLVR), where the reward is typically binary:
 r ​ ( x , y ) = 1 r(x,y)=1 if the solution y is correct and 0 otherwise
 (DeepSeek-AI et al., 2025 ) . This binary, outcome-level reward
makes credit assignment particularly challenging, as the model must
determine which tokens contributed to a correct or incorrect answer.

 2.2 Group Relative Policy
Optimization

GRPO (Shao et al., 2024 ) generates G G rollouts
 { y 1 , … , y G } \left\{y_{1},\ \ldots,\ y_{G}\right\} for each prompt x x and
computes advantages by normalizing rewards within each group:
 A i ^ = r i − μ σ \widehat{A_{i}}\ =\ \frac{r_{i}\ -\ \mu}{\sigma} . This
eliminates the need for a separate critic network, reducing memory
requirements and avoiding the instability of value function estimation.

GRPO applies PPO-style ratio clipping (Schulman et al., 2017 ) to
stabilize updates:
 L G ​ R ​ P ​ O = − E ​ [ min ⁡ ( ρ t ​ A , c ​ l ​ i ​ p ​ ( ρ t , 1 − ε , 1 + ε ) ​ A ) ] L_{GRPO}\ =\ -E\left[\min\left(\rho_{t}A,\ clip\left(\rho_{t},\ 1-\varepsilon,\ 1+\varepsilon\right)A\right)\right] ,
where
 ρ t = π θ ​ ( a t | s t ) π o ​ l ​ d ​ ( a t | s t ) \rho_{t}\ =\ \frac{\pi_{\theta}\left(a_{t}|s_{t}\right)}{\pi_{old}\left(a_{t}|s_{t}\right)} 
is the importance ratio. Training alternates between generating rollouts
from the current policy and performing gradient updates on the clipped
objective.

 2.3 The Cliff Problem

Under binary reward with G rollouts, the advantage for a prompt x takes
one of three forms: (i) all rollouts succeed, giving zero advantages —
the model has already learned this prompt; (ii) mixed results, giving
positive advantages for successes and negative for failures — the
standard learning regime; or (iii) all rollouts fail, giving zero
advantages — the model receives no gradient despite urgently needing
signal.

Case (iii) — commonly called a “cliff” (Zhang et al., 2025b ) — is the critical failure
mode: the hardest problems, which represent the frontier of the model’s
capability, receive essentially no gradient
 (Le et al., 2025 ) . The model can only learn from prompts
at intermediate difficulty — hard enough that some rollouts fail
(providing contrast for advantage estimation) but easy enough that some
succeed. The cliff boundary can only advance indirectly, as learning on nearby prompts
must transfer to cliff prompts through weight sharing alone — there is no
direct gradient signal on the cliffs themselves.

 2.4 Approaches to the Cliff
Problem

A growing body of work has identified the cliff problem as a critical
bottleneck in RL for reasoning and proposed diverse strategies to
address it. These approaches introduce substantial additional complexity
— new hyperparameters, auxiliary models, replay infrastructure, or
modifications to the training loop — to work around the zero-gradient
problem. By contrast, HDPO requires only a single additional forward
pass with ground truth appended to the prompt and a standard JSD loss:
no curriculum scheduler, no replay buffer, no process reward model, no
scaffolding heuristics. We organize the existing approaches into several
categories to illustrate this complexity.

 Curriculum and filtering . VCRL (Jiang et al., 2025 ) 
observes that group reward variance is a natural proxy for prompt
difficulty and dynamically schedules sampling to focus on the
high-variance frontier. DAPO (Yu et al., 2025 ) filters
out zero-variance prompts entirely, oversampling until informative prompts
are found. While
effective, these methods avoid cliff prompts rather than learning from
them — the hardest problems are simply skipped,
leaving the cliff boundary to advance only through indirect transfer
from nearby solved prompts.

 Scaffolded and augmented generation . Scaf-GRPO
 (Zhang et al., 2025b ) diagnoses learning stagnation on cliff
prompts and injects tiered in-prompt hints — from abstract concepts to
concrete steps — that fade as the model improves. HINT
 (Wang et al., 2025 ) provides targeted guidance to help
ineffective rollouts navigate toward correct solutions. EvoCoT
 (Liu et al., 2025 ) constrains the exploration space by
self-generating and verifying chain-of-thought trajectories, then
gradually expands by shortening CoT steps. These methods require
designing or learning a hint generation strategy, deciding when to
inject and fade hints, and managing the distributional gap between
scaffolded and unscaffolded generations — a non-trivial engineering
and research effort per domain.

 Experience replay . Retrospective Replay
 (Dou et al., 2025 ) stores early exploratory trajectories
and revisits them later when the model has grown capable. RLEP
 (Zhang et al., 2025a ) collects verified correct trajectories
and blends them into mini-batches. Replay methods introduce a buffer
management system with its own hyperparameters (buffer size, sampling
strategy, staleness thresholds) and face a fundamental tension: replayed
trajectories are off-policy and may become increasingly stale as
training progresses, potentially teaching the model to imitate outdated
reasoning patterns.

 Advantage shaping and process rewards . Le et al.
 (Le et al., 2025 ) tackle zero-variance prompts by scaling
each token’s gradient in proportion to its entropy, reclaiming learning
signal from otherwise wasted prompts. PRIME
 (Cui et al., 2025 ) enables online process reward model
updates using only policy rollouts and outcome labels. Lightman et al.
 (Lightman et al., 2024 ) demonstrate that step-level
verification rewards substantially improve reasoning. However, process
reward approaches require training and maintaining a separate reward
model — itself a significant undertaking — and shaped rewards may
not perfectly align with the outcome objective.

 Interleaved RL and distillation . ReLIFT
 (Ma et al., 2025 ) finds that RL excels on easier
questions while SFT is crucial for the hardest ones, and interleaves RL
with targeted online SFT on questions the current policy cannot solve.
ReLIFT is conceptually closest to HDPO but is considerably more complex:
it requires a two-phase training loop with switching logic, an external
source of high-quality solutions for the SFT phase, and careful tuning
of the interleaving schedule. HDPO achieves the same goal — learning
signal on cliff prompts — with a single unified objective, using the
model’s own privileged rollouts rather than external solutions, and with
theoretical guarantees (bounded realizability gap, optimal target
characterization) that ReLIFT lacks.

In summary, existing approaches to the cliff problem introduce
substantial machinery — curriculum schedulers, hint generators, replay
buffers, process reward models, or multi-phase training loops — to
work around the zero-gradient problem. HDPO’s mechanism is remarkably
simple by comparison: append ground truth to the prompt, generate,
filter for correctness, and distill via JSD. This simplicity is not a
limitation but a feature: it arises from the observation that the model
can solve cliff prompts when given privileged information — the same
weights that fail unprompted succeed with ground-truth context, providing
a naturally bounded distillation target. HDPO is also
complementary to these approaches — curriculum methods could
prioritize near-cliff prompts, and process rewards could densify signal
on non-cliff prompts — but it achieves its core effect with minimal
additional complexity.

 2.5 Knowledge Distillation for
Reasoning

Knowledge distillation (Hinton et al., 2015 ) transfers
knowledge from a teacher to a student by minimizing the divergence
between their output distributions. For autoregressive language models,
a key challenge is distribution mismatch: the student is trained on
teacher-generated or fixed sequences, but at inference must generate
from its own distribution. Agarwal et al.
 (Agarwal et al., 2024 ) address this with Generalized Knowledge
Distillation (GKD), which trains the student on its own generated
outputs while receiving teacher feedback. GKD offers flexibility in
divergence choice (forward KL, reverse KL, JSD).

 Self-distillation for reasoning . Zhao et al.
 (Zhao et al., 2026 ) propose on-policy self-distillation
(OPSD), where a single LLM serves as both teacher and student under
different informational contexts — the teacher conditions on
privileged information such as verified reasoning traces or ground-truth
solutions, while the student sees only the problem. Training minimizes
per-token divergence over the student’s own on-policy rollouts. Hübotter
et al. (Hübotter et al., 2026 ) introduce Self-Distillation
Policy Optimization (SDPO), which converts rich textual feedback
(runtime errors, judge evaluations) into a dense learning signal by
distilling the model’s feedback-conditioned predictions back into the
unconditional policy, addressing the credit-assignment bottleneck of
scalar-only reward.

 Unified KD and RL frameworks. Several works jointly optimize
distillation and RL objectives. KDRL (Xu et al., 2025 ) 
simultaneously minimizes reverse KL divergence between student and
teacher while maximizing expected reward, finding that the combination
outperforms either objective alone. RLAD
 (Zhang et al., 2026 ) replaces the standard importance ratio
with a geometric mixture of the old policy and a teacher model,
embedding the teacher directly into the RL policy update. G-OPD
 (Yang et al., 2026 ) shows theoretically that standard
on-policy distillation is a special case of dense KL-constrained RL and
introduces reward extrapolation to enable students to surpass teacher
performance.

 Diversity preservation . Li et al.
 (Li et al., 2025 ) identify the diversity collapse
problem: RLVR fine-tuning improves pass@1 but degrades pass@k as the
policy concentrates around a single mode. DPH-RL replaces mode-seeking
divergences (reverse KL) with mass-covering alternatives (forward KL,
JSD) that continuously reference the initial policy to maintain broad
solution coverage. This insight directly motivates HDPO’s use of JSD for
the distillation loss: JSD provides mode-covering signal that expands
the student’s support without collapsing to a single mode.

 Privileged information . The concept of learning using
privileged information (LUPI) (Vapnik and Vashist, 2009 ) provides a
theoretical framework for training with auxiliary information available
only at training time. Penaloza et al. (Penaloza et al., 2026 ) 
introduce pi-Distill, a joint teacher-student objective for privileged
information distillation. HDPO instantiates the LUPI framework
specifically for RL: ground truth serves as privileged information, the
model itself serves as teacher, JSD (Lin, 1991 ) 
provides the distillation mechanism, and distillation targets only cliff
prompts where the RL gradient vanishes.

 3 Method: Hybrid Distillation Policy
Optimization

 3.1 Motivation

The central insight of HDPO is that for cliff prompts, we can construct
a teacher signal without any external model or human annotation — by
giving the model itself access to the answer. When conditioned on
ground-truth information (e.g., the solution to a math problem), even a
small model can generate correct reasoning traces with high probability.
The key properties of this construction are: (i) the teacher and student
share the same weights, differing only in their input context; (ii) the
 R = 1 R=1 filter selects only correct teacher trajectories; and (iii)
distillation targets only cliff prompts where the RL gradient is zero.

Property (i) is critical: because teacher and student are the same
model, the realizability gap — the KL divergence between their output
distributions — is bounded by a function of the privileged information
alone (Proposition 1). This is strictly tighter than the bound for any
cross-model teacher, which incurs an additional model-mismatch term that
same-model distillation eliminates entirely. Property (ii) ensures we
distill from the optimal target, not just any correct trajectory.
Property (iii) restricts distillation to cliff prompts, where the RL
gradient is zero and thus cannot provide signal on its own.

 3.2 Formal Definition

Formal Definition.

The HDPO training objective combines
 ℒ G ​ R ​ P ​ O \mathcal{L}_{GRPO} with a ℒ J ​ S ​ D \mathcal{L}_{JSD} distillation term on
cliff prompts:

 ℒ H ​ D ​ P ​ O ​ ( θ ) = ℒ G ​ R ​ P ​ O ​ ( θ ) + λ ⋅ ℒ J ​ S ​ D ​ ( θ ) \mathcal{L}_{HDPO}(\theta)\ =\ \mathcal{L}_{GRPO}(\theta)\ +\ \lambda\ \cdot\ \mathcal{L}_{JSD}(\theta) 

where ℒ J ​ S ​ D \mathcal{L}_{JSD} is the token-averaged JSD over filtered
teacher trajectories on cliff prompts:

 ℒ J ​ S ​ D ( θ ) = 1 N t ​ o ​ k ∑ ( x , y ) ∈ 𝒯 ∑ t = 1 | y | J S D k ( π T ( ⋅ ∣ y < t ) ∥ π θ ( ⋅ ∣ y < t ) ) \mathcal{L}_{JSD}(\theta)\ =\ \frac{1}{N_{tok}}\sum_{(x,\ y)\mathcal{\in T}}{\sum_{t\ =\ 1}^{|y|}{{JSD}_{k}\left(\pi_{T}\left(\ \cdot\ \mid\ y_{<\ t}\right)\ \parallel\ \pi_{\theta}\left(\ \cdot\ \mid\ y_{<\ t}\right)\right)}} 

The distillation set 𝒯 \mathcal{T} is constructed via two levels of
filtering. First, identify cliff prompts — prompts where all K K 
standard rollouts failed:

 𝒞 = { x ∈ ℬ : ∑ k R ​ ( x , y ( k ) ) = 0 } \mathcal{C}=\{x\in\mathcal{B}:\textstyle\sum_{k}R(x,y^{(k)})=0\} 

Second, for each x ∈ 𝒞 x\ \in\ \mathcal{C} , generate privileged rollouts
 y ¯ j ∼ π θ ( ⋅ | x ⊕ y ∗ ) x \bar{y}_{j}\ \sim\ \pi_{\theta}(\cdot\ |\ x\ \oplus\ y*)_{x} by
injecting ground truth y ∗ y^{*} into the prompt, and retain only
correct trajectories:

 𝒯 = { ( x , y ¯ ) : x ∈ 𝒞 , y ¯ ∼ π θ ( ⋅ ∣ x ⊕ y ∗ ) , R ( x , y ¯ ) = 1 } \mathcal{T}=\{(x,\ \bar{y})\ :\ x\in\mathcal{C},\ \bar{y}\sim\pi_{\theta}(\cdot\mid x\oplus y^{*}),\ R(x,\ \bar{y})=1\} 

That is, 𝒯 \mathcal{T} contains (prompt, trajectory) pairs where (a)
the prompt is a cliff — all standard rollouts scored zero — and (b)
the privileged model, conditioned on ground truth, generated a correct
solution.
 N t ​ o ​ k = ∑ ( x , y ¯ ) ∈ 𝒯 | y ¯ | N_{tok}\ =\ \sum_{\left(x,\ \overline{y}\right)\mathcal{\in T}}\left|\overline{y}\right| 
is the total distillation token count (computed globally across all
data-parallel ranks). π T \pi_{T} and π θ \pi_{\theta} share the same
weights; they differ only in their input (privileged vs. unprivileged).
In practice, J ​ S ​ D k {JSD}_{k} is approximated using the
teacher's top- k k ( k = 64 k\ =\ 64 ) logits,
renormalized. The tail correction for student mass outside the top- k k 
support is exact: P r ​ e ​ s ​ t ⋅ l ​ n ​ 2 P_{rest}\ \cdot\ ln\ 2 .

 3.3 Algorithm

 Algorithm 1 HDPO Training 

 0:  Policy π θ \pi_{\theta} , prompt set 𝒳 \mathcal{X} , ground truth { y ∗ } \{y^{*}\} ,
reward R R , learning rate α \alpha , distillation weight λ \lambda ,
rollouts per prompt K K 

 1:   for step = 1 , … , N =1,\ldots,N do 

 2:    // Standard GRPO 

 3:   Sample prompt batch B ⊂ 𝒳 B\subset\mathcal{X} 

 4:    for all x ∈ B x\in B do 

 5:    Generate K K rollouts y ( k ) ∼ π θ ( ⋅ ∣ x ) y^{(k)}\sim\pi_{\theta}(\cdot\mid x) 

 6:    end for 

 7:   Score: A ^ i = ( r i − μ ) / σ \hat{A}_{i}=(r_{i}-\mu)/\sigma 

 8:   Compute ℒ GRPO \mathcal{L}_{\mathrm{GRPO}} via clipped policy gradient with leave-one-out advantages

 9:    // Privileged self-distillation on cliff prompts 

 10:   Identify cliffs: 𝒞 = { x ∈ B : ∑ k r ( k ) = 0 } \mathcal{C}=\{x\in B:\textstyle\sum_{k}r^{(k)}=0\} 

 11:    for all x ∈ 𝒞 x\in\mathcal{C} do 

 12:    Generate y ¯ ( j ) ∼ π θ ( ⋅ ∣ x ⊕ y ∗ ) \bar{y}^{(j)}\sim\pi_{\theta}(\cdot\mid x\oplus y^{*}) 

 13:    end for 

 14:   Filter: 𝒯 = { ( x , y ¯ ) : R ​ ( x , y ¯ ) = 1 } \mathcal{T}=\{(x,\bar{y}):R(x,\bar{y})=1\} 

 15:   Compute ℒ JSD = 1 N tok ∑ ( x , y ¯ ) ∈ 𝒯 ∑ t = 1 | y ¯ | JSD k ( π T ( ⋅ | y ¯ < t ) ∥ π θ ( ⋅ | y ¯ < t ) ) \mathcal{L}_{\mathrm{JSD}}=\frac{1}{N_{\mathrm{tok}}}\sum_{(x,\bar{y})\in\mathcal{T}}\sum_{t=1}^{|\bar{y}|}\mathrm{JSD}_{k}(\pi_{T}(\cdot|\bar{y}_{<t})\,\|\,\pi_{\theta}(\cdot|\bar{y}_{<t})) 

 16:    // Update 

 17:    θ ← θ − α ​ ∇ θ ( ℒ GRPO ​ ( θ ) + λ ⋅ ℒ JSD ​ ( θ ) ) \theta\leftarrow\theta-\alpha\nabla_{\theta}\bigl(\mathcal{L}_{\mathrm{GRPO}}(\theta)+\lambda\cdot\mathcal{L}_{\mathrm{JSD}}(\theta)\bigr) 

 18:   end for 

 3.4 Theoretical Analysis

We argue that HDPO’s effectiveness rests on two properties: on the
correct teacher trajectories, the distance between the student and
teacher distributions is bounded (unlike cross-model distillation), and
the R = 1 R\ =\ 1 filter ensures the teacher distribution the student
moves toward corresponds to the KL-regularized RL optimal policy
(Proposition 2).
Moreover,
this bound is strictly tighter than what any cross-model teacher can
achieve: because teacher and student share the same function, the only
source of distributional divergence is the privileged information
itself. A cross-model teacher introduces an additional model-mismatch
term that same-model distillation eliminates entirely (Proposition 1).

 Proposition 1 (Realizability Gap Comparison). Let P T P_{T} 
 and P S P_{S} denote the per-position output distributions
of a language model θ \theta when prompted with and without
ground truth g g , respectively. 

 D K ​ L ​ ( P T ∥ P S ) ≤ L θ 2 ⋅ Δ ​ ( g ) 2 2 D_{KL}(P_{T}\parallel P_{S})\leq\frac{L_{\theta}^{2}\cdot\Delta(g)^{2}}{2} 

 (i) Same-model bound. where L θ L_{\theta} is the
local Lipschitz constant of the model’s logit function on bounded inputs
and Δ ​ ( g ) \Delta(g) is the input-space distance attributable to 
 g g . 

 (ii) Cross-model comparison. For cross-model distillation
with a separate teacher model ϕ \phi receiving the same
privileged input c T c_{T} , the realizability gap satisfies: 

 D K ​ L ( P ϕ ( ⋅ | c T ) ∥ P θ ( ⋅ | c S ) ) ≤ ( L θ ⋅ Δ ​ ( g ) + ‖ f ϕ ​ ( c T ) − f θ ​ ( c T ) ‖ ∞ ) 2 2 D_{KL}(P_{\phi}(\cdot|c_{T})\parallel P_{\theta}(\cdot|c_{S}))\leq\frac{\left(L_{\theta}\cdot\Delta(g)+\left\|f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{T}\right)\right\|_{\infty}\right)^{2}}{2} 

 The cross-model bound contains an additive model-mismatch term 
 ‖ f ϕ ​ ( c T ) − f θ ​ ( c T ) ‖ ∞ \left\|f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{T}\right)\right\|_{\infty} 
 that is absent from the same-model bound. This term vanishes if
and only if the models produce identical logits on the privileged
input. 

Proof sketch.

Both bounds follow from Lemma 1 (Appendix A),
which shows
 D K ​ L ≤ ‖ δ ‖ ∞ 2 2 D_{KL}\leq\frac{{\left\|\delta\right\|_{\infty}}^{2}}{2} for any
logit perturbation δ \delta . For (i): since teacher and student
evaluate the same function f θ f_{\theta} on different inputs, the logit
difference is bounded by L θ ⋅ Δ ​ ( g ) L_{\theta}\cdot\Delta(g) via Lipschitz
continuity. For (ii): the logit difference decomposes via the triangle
inequality into a model-mismatch term plus the same input-perturbation
term from (i) (Appendix A).

Interpretation.

The proposition provides a direct comparison
between same-model and cross-model distillation. Both bounds share the
same Δ ​ ( g ) \Delta(g) term — the cost of incorporating privileged
information — but the cross-model bound pays an additional price for
model mismatch. With a drifting teacher (where the teacher
shares the student’s current weights), the model-mismatch term is
identically zero throughout training. With a frozen teacher 
(initialized at the start of training), a non-zero model-mismatch term
appears, yielding looser guarantees that approach the cross-model regime.

Relation to prior work.

The bound follows from standard
Lipschitz continuity and logit perturbation arguments. Lemma 1
(Appendix A) is a known bound on KL divergence under softmax logit
perturbation. The contribution of Proposition 1 is not the bound itself
but the comparison: it makes explicit that same-model privileged
distillation eliminates the model-mismatch term that cross-model
distillation cannot avoid.

 Proposition 2 ( R = 1 R{=}1 Filtering Yields the RL-Optimal
Policy). Consider the KL-regularized RL objective: 

 max π ⁡ 𝔼 τ ∼ π ​ [ R ​ ( τ ) ] − β ⋅ K ​ L ​ ( π ∥ π r ​ e ​ f ) \max_{\pi}\ \mathbb{E}_{\tau\sim\pi}[R(\tau)]\ -\ \beta\ \cdot\ KL(\pi\ \parallel\ \pi_{ref}) 

 For binary reward R ∈ { 0 , 1 } R\in\{0,1\} , the unique optimizer
is π ∗ ​ ( x ) ∝ π r ​ e ​ f ​ ( x ) \pi^{*}(x)\propto\pi_{ref}(x) e ​ x ​ p ​ ( R ​ ( x ) / β ) exp(R(x)/\beta) . In
the limit β → 0 + \beta\rightarrow 0^{+} , π ∗ \pi^{*} converges
to π r ​ e ​ f ( ⋅ | R = 1 ) \pi_{ref}(\cdot|R=1) — the reference policy
conditioned on correctness — and R = 1 R=1 rejection sampling
from π r ​ e ​ f \pi_{ref} recovers it exactly. 

Proof sketch.

The unique solution takes the form  (Rafailov et al., 2023 ) :

 π ∗ ​ ( τ ) = π r ​ e ​ f ​ ( τ ) ⋅ exp ⁡ ( R ​ ( τ ) β ) Z ​ ( β ) \pi^{*}(\tau)=\ \pi_{ref}(\tau)\cdot\frac{\exp\left(\frac{R(\tau)}{\beta}\right)}{Z(\beta)} 

where Z ​ ( β ) Z(\beta) is the partition function. For binary reward
 R ​ ( τ ) ∈ { 0 , 1 } R(\tau)\in\{0,1\} , this reweights correct trajectories by
 e ​ x ​ p ​ ( 1 / β ) exp(1/\beta) relative to incorrect ones. In the hard-threshold limit
 β → 0 + \beta\rightarrow 0^{+} , incorrect trajectories receive zero weight and
the optimal policy concentrates entirely on correct trajectories
(Appendix B):

 π ∗ ​ ( τ ) → π r ​ e ​ f ​ ( τ | R ​ ( τ ) = 1 ) = π r ​ e ​ f ​ ( τ ) ⋅ 1 ​ [ R ​ ( τ ) = 1 ] P π r ​ e ​ f ​ ( R = 1 ) \pi^{*}(\tau)\ \rightarrow\ \pi_{ref}(\tau\ |\ R(\tau)\ =\ 1)\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ 1[R(\tau)\ =\ 1]}{P_{\pi_{ref}}(R\ =\ 1)} 

Connection to HDPO.

The KL-regularized objective above is the
standard RL formulation: maximize reward while staying close to a
reference π r ​ e ​ f \pi_{ref} . Proposition 2 shows that in the
 β → 0 \beta\ \rightarrow\ 0 limit with binary reward, the optimal policy
is π r ​ e ​ f ( ⋅ | R = 1 ) \pi_{ref}(\cdot\ |\ R\ =\ 1) — the reference distribution
conditioned on correctness — and that R = 1 R=1 rejection sampling
recovers it exactly. On non-cliff prompts, GRPO's policy
gradient optimizes this objective directly and provides learning signal.
On cliff prompts, however, all K K rollouts score zero:
 P π _ θ ( ⋅ | p r o m p t ) ​ ( R = 1 ) ≈ 0 P_{\pi\_\theta(\cdot|prompt)}(R\ =\ 1)\ \approx\ 0 . The
reference has negligible support on correct trajectories, so
 π r ​ e ​ f ( ⋅ | R = 1 ) \pi_{ref}(\cdot\ |\ R\ =\ 1) is inaccessible — rejection
sampling produces no accepted samples, and the RL gradient vanishes.

HDPO resolves this by replacing π r ​ e ​ f \pi_{ref} with a proxy whose optimal
solution has non-degenerate support on correct trajectories. Injecting
ground truth g g into the prompt yields
 π θ ( ⋅ | p r o m p t , g ) \pi_{\theta}(\cdot\ |\ prompt,\ g) , which shares the
student's weights but has substantially higher
 P ​ ( R = 1 ) P(R\ =\ 1) . The KL-regularized objective with this proxy
reference,

 max π 𝔼 [ R ( τ ) ] − β ⋅ K L ( π ∥ π θ ( ⋅ | p r o m p t , g ) ) \max\ _{\pi}\mathbb{\ E[}R(\tau)]\ -\ \beta\ \cdot\ KL(\pi\ \parallel\ \pi_{\theta}(\cdot\ |\ prompt,\ g)) 

has a well-defined optimal solution
 π θ ( ⋅ | p r o m p t , g , R = 1 ) \pi_{\theta}(\cdot\ |\ prompt,\ g,\ R\ =\ 1) with non-empty
support on correct trajectories, precisely where the original
objective's solution
 π θ ( ⋅ | p r o m p t , R = 1 ) \pi_{\theta}(\cdot\ |\ prompt,\ R\ =\ 1) had none. Ground truth
injection expands the set of trajectories τ \tau where
 R ​ ( τ ) = 1 R(\tau)\ =\ 1 , transforming an objective with a degenerate
solution into one with a substantive target. R = 1 R=1 filtering on this
proxy has a well-defined optimal solution (Proposition 2). HDPO
approximates it by generating from π θ ( ⋅ ∣ x , y ∗ ) \pi_{\theta}(\cdot\mid x,y^{*}) and
retaining only R=1 completions, yielding a finite-sample estimate of
 π θ ( ⋅ ∣ x , y ∗ , R = 1 ) \pi_{\theta}(\cdot\mid x,y^{*},R{=}1) . This departs from the theoretical target
in two ways: the proxy reference π priv \pi_{\mathrm{priv}} replaces π ref \pi_{\mathrm{ref}} (the gap
bounded by Proposition 1), and the privileged pass rate can be substantially
below 1 on hard problems, introducing sampling noise. Proposition 1 then
bounds the gap between this proxy target and the ideal target
 π θ ( ⋅ | p r o m p t , R = 1 ) \pi_{\theta}(\cdot\ |\ prompt,\ R\ =\ 1) : because teacher and
student share weights, the only source of divergence is the privileged
information itself, with the model-mismatch term identically zero for a
drifting teacher.

 π \pi -Distill (Penaloza et al., 2026 ) independently arrives at the same objective
structure: its teacher objective
 J T ​ e ​ a ​ c ​ h ​ e ​ r = 𝔼 [ R ] − β ⋅ D K ​ L ( π T ( ⋅ | x , I ) ∥ s g [ π S ( ⋅ | x ) ] ) J_{Teacher}\mathbb{\ =\ E[}R]\ -\ \beta\ \cdot\ D_{KL}(\pi_{T}(\cdot|x,\ I)\ \parallel\ sg[\pi_{S}(\cdot|x)]) 
takes exactly this form with π r ​ e ​ f \pi_{ref} = sg[ π S \pi_{S} ]. Where
 π \pi -Distill optimizes this objective via gradient ascent at finite
 β \beta across all prompts, HDPO implements the
 β → 0 \beta\ \rightarrow\ 0 limit mechanically via hard R=1 filtering,
activating only on cliff prompts where gradient-based optimization of
the objective fails.

 4 Experiments

 4.1 Experimental Setup

We evaluate HDPO using OpenMathInstruct-2
 (Toshniwal et al., 2024 ) , a large-scale dataset of mathematical
problems sourced from the MATH competition benchmark
 (Hendrycks et al., 2021 ) and GSM8K. We use
Qwen2.5-Math-1.5B-Instruct (Yang et al., 2024 ) as the base
model. The dataset is split into 95% training and 5% validation (seed
42), yielding 2048 validation problems evaluated with binary rewards
from hf_math_verify. The base RL algorithm is GRPO with G=16 rollouts
per prompt, leave-one-out advantages, and PPO-style ratio clipping
( ε = 0.2 \varepsilon{=}0.2 ). We train for 2000 steps with AdamW
 (Loshchilov and Hutter, 2019 ) at learning rate 1 ​ e − 6 1\mathrm{e}{-}6 with linear warmup.
Generation uses vLLM (Kwon et al., 2023 ) at temperature
1.0. Full experimental details are provided in Appendix C.

We compare four HDPO configurations against the GRPO baseline, varying
the teacher type (frozen at initialization vs. drifting with current
policy) and the distillation weight ( λ \lambda ∈ \in {0.01, 0.1}). All
HDPO variants use top-k=64 JSD with global token-count normalization. We
report best pass@1, pass@4, and pass@8 across training, evaluated every
10 steps with 2048 samples per evaluation.

 4.2 Results

 Table 1: Best pass@ k k on OpenMathInstruct-2 validation (2048 samples,
Qwen2.5-Math-1.5B-Instruct) across 2000 training steps on 8 × \times H200 GPUs.
All HDPO variants use JSD with global token-count normalization.
Bold indicates best in column. Results on 8 × \times H100 GPUs are provided in Appendix  D and show consistent trends with minor quantitative variation. 

 Method 
 pass@1 
 pass@4 
 pass@8 

 GRPO Baseline 
 0.6519 
 0.7749 
 0.8228 

 HDPO (frozen, λ \lambda =0.01) 
 0.6519 
 0.7812 
 0.8218 

 HDPO (frozen, λ \lambda =0.1) 
 0.6304 
 0.7812 
 0.8398 

 HDPO (drifting, λ \lambda =0.01) 
 0.6514 
 0.7861 
 0.8271 

 HDPO (drifting, λ \lambda =0.1) 
 0.6294 
 0.7856 
 0.8364 

The drifting teacher at λ = 0.01 \lambda\ =\ 0.01 achieves the highest
pass@4 (0.7861, +1.1% over baseline) and strong pass@8 (0.8271, +0.4%) while essentially
maintaining pass@1 (0.6514 vs 0.6519). This is our primary
result — HDPO broadens the model’s coverage without sacrificing
greedy accuracy. We note that the magnitude of improvement at λ = 0.01 \lambda{=}0.01 varies across hardware (see Appendix  D ); the effect is consistent in direction but modest in size relative to run-to-run variance.

Increasing λ \lambda from 0.01 to 0.1 trades pass@1 for pass@8. Both
teacher types at λ = 0.1 \lambda\ =\ 0.1 achieve ~0.84
pass@8 (+1.4–1.7% over baseline) but pass@1 drops by
~2.3–2.8%. λ \lambda directly controls the
exploration–exploitation trade-off in the output distribution. The λ = 0.1 \lambda{=}0.1 improvements on pass@8 are the most robust finding, reproducing consistently across hardware configurations.

The drifting teacher advantage is most pronounced at low λ \lambda ;
at λ = 0.1 \lambda\ =\ 0.1 , the advantage narrows as the strong
distillation signal dominates.

 5 Discussion

The results demonstrate that HDPO provides a consistent mechanism for
broadening the model’s solution distribution on mathematical reasoning
tasks. At λ = 0.01 \lambda{=}0.01 , HDPO improves pass@4 and pass@8 while essentially maintaining pass@1, though the magnitude of improvement is modest and varies across hardware (Appendix  D ). At λ = 0.1 \lambda{=}0.1 , the coverage gains are larger and more robust, at the cost of pass@1. This is the core value proposition: by providing
learning signal on cliff prompts where GRPO cannot, HDPO expands the
model’s coverage, with λ \lambda controlling how aggressively it trades greedy accuracy for coverage.

The λ \lambda parameter provides explicit control over the
exploration–exploitation tradeoff. At λ \lambda =0.01, the
distillation signal is a gentle nudge that broadens coverage; at
 λ \lambda =0.1, it becomes the dominant training signal on cliff
prompts, significantly improving pass@8 but at the cost of pass@1. At our model scale (1.5B), we observe a consistent tradeoff between
pass@1 and pass@ k k as λ \lambda increases, though it remains an open
question whether this tradeoff persists at larger scales.

We hypothesize that this tradeoff reflects model capacity interacting
with mode-covering distillation. When the privileged model solves a
problem via multiple distinct strategies, the JSD loss trains the
student to place mass on all of them. For a small model, limited
capacity means these modes compete for the same parameters: the model
cannot maintain multiple fully coherent reasoning strategies
simultaneously. The result is a flatter distribution where no single
strategy dominates cleanly, degrading greedy accuracy (pass@1), while
the broader support means additional samples discover genuinely
different approaches (improving pass@k). We hypothesize that the ideal output distribution
is not uniform but concentrated: a single dominant mode that greedy
decoding reliably recovers, with smaller secondary modes accessible
through repeated sampling. Pure RL naturally produces this shape through
mode-seeking dynamics but cannot discover strategies where all rollouts
fail. This analysis motivates the expand-then-sharpen curriculum
described in Section 6: distillation first broadens strategy support on
cliff prompts, then RL restores a dominant mode while preserving
secondary strategies as a long tail.

The drifting teacher — which shares the current policy’s weights at
each step — outperforms the frozen teacher (initialized at the start
of training) at low λ \lambda . At λ = 0.1 \lambda=0.1 , the advantage
narrows as the strong distillation signal dominates.

The frozen teacher at λ \lambda =0.1 achieves the single highest pass@8
(0.8398), suggesting that the initial model’s distribution may be more
diverse — not yet shaped by RL’s mode-seeking dynamics — at the cost
of a larger realizability gap. These results suggest a tradeoff between signal diversity and
realizability in teacher choice, though confirming this as a general
principle requires broader evaluation across scales and datasets.

 6 Limitations

Our results are demonstrated on a single model scale (1.5B parameters)
and a single dataset (OpenMathInstruct-2, sourced from MATH
 (Hendrycks et al., 2021 ) ). While the theoretical analysis (Proposition
1, Proposition 2) is scale-agnostic, the empirical magnitude of HDPO’s
improvements may vary with model capacity, dataset difficulty, and
reward function design. Larger models may have fewer cliff prompts
(higher baseline success rates), potentially reducing HDPO’s marginal
benefit, or may benefit more from the expanded coverage.

HDPO introduces computational overhead: generating and filtering
privileged rollouts, computing top-k teacher logits, and the additional
forward pass for the JSD loss. In our implementation, privileged
generation is performed by the same vLLM instance used for standard
rollouts, amortizing some of this cost. The overhead is proportional to
the number of cliff prompts per step.

A natural extension is to re-inject previously cliff prompts — now
solvable after distillation — back into the RL training distribution
for mode-sharpening. Crucially, re-injection should be delayed: by
allowing continued training on other prompts before revisiting former
cliffs, we test whether the distilled strategies are durably encoded
rather than transiently accessible. This expand-then-sharpen cycle could
systematically convert cliffs into solved prompts while maintaining the
benefits of RL’s mode-seeking dynamics, achieving both high pass@k
(broad coverage) and high pass@1 (sharp modes). We leave this
curriculum-learning extension to future work.

 7 Conclusion

We have introduced HDPO, a hybrid training objective that augments
reinforcement learning with privileged self-distillation to address the
cliff problem in mathematical reasoning. By leveraging ground truth as
privileged information and the model’s own weights as the teacher, HDPO
provides bounded, non-zero gradients on prompts where the standard RL
gradient vanishes. Our theoretical analysis shows that same-model
privileged distillation achieves a strictly tighter realizability gap
than cross-model distillation, with the gap depending only on the
model’s Lipschitz constant and the information content of the ground
truth (Proposition 1), and that R=1 filtering recovers the optimal
KL-regularized RL policy (Proposition 2).

Experiments on OpenMathInstruct-2 with Qwen2.5-Math-1.5B-Instruct show
that HDPO consistently improves p ​ a ​ s ​ s ​ @ ​ 4 pass@4 and p ​ a ​ s ​ s ​ @ ​ 8 pass@8 , with the distillation weight λ \lambda 
providing direct control over the exploration–exploitation tradeoff.
At λ = 0.1 \lambda{=}0.1 , coverage improvements (pass@8 + 1.4 +1.4 – 1.7 % 1.7\% ) reproduce robustly across hardware; at λ = 0.01 \lambda{=}0.01 , improvements are consistent in direction but smaller in magnitude and more sensitive to run-to-run variance.

Looking forward, the expand-then-sharpen paradigm — using HDPO to
broaden coverage, then RL to sharpen modes on previously unsolvable
prompts — suggests a curriculum that could progressively reduce the fraction
of cliff prompts. Combined with evaluation on larger model scales
and diverse reasoning benchmarks, this direction could establish
privileged self-distillation as a standard component of RL training for
language models.

 Appendix A Realizability Gap Bound — Full
Proof

 A.1 Setup

Let f θ f_{\theta} denote a language model with parameters θ \theta 
and vocabulary 𝒱 \mathcal{V} . For context c c (a sequence of token
embeddings), the model produces logits
 z = f θ ​ ( c ) ∈ ℝ | 𝒱 | z=f_{\theta}(c)\in\mathbb{R}^{\mathcal{|V|}} and output
distribution P θ ( ⋅ | c ) = s o f t m a x ( f θ ( c ) ) P_{\theta}(\cdot|c)=softmax(f_{\theta}(c)) .

For prompt x x , ground truth g g , and prefix y < t y_{<t} , define:

 • 

 Teacher : P T = P θ ( ⋅ ∣ c T ) P_{T}=P_{\theta}(\cdot\mid c_{T}) where
 c T = [ x ; g ; y < t ] c_{T}=[x;\,g;\,y_{<t}] 

 • 

 Student : P S = P θ ( ⋅ ∣ c S ) P_{S}=P_{\theta}(\cdot\mid c_{S}) where
 c S = [ x ; y < t ] c_{S}=[x;\,y_{<t}] 

The realizability gap at position t t is
 D K ​ L ​ ( P T ∥ P S ) D_{KL}(P_{T}\parallel P_{S}) : the KL divergence between the teacher
and student distributions. A small realizability gap means the student
can feasibly match the teacher’s distribution without large parameter
updates — the optimization target is within reach.

 A.2 Lemma 1 (KL Divergence under Logit
Perturbation)

This is a standard bound in information theory; we include the proof for
completeness.

 Statement. For any z ∈ ℝ | 𝒱 | z\in\mathbb{R}^{\mathcal{|V|}} 
 and δ ∈ ℝ | 𝒱 | \delta\in\mathbb{R}^{\mathcal{|V|}} : 

 D K ​ L ​ ( s ​ o ​ f ​ t ​ m ​ a ​ x ​ ( z ) ∥ s ​ o ​ f ​ t ​ m ​ a ​ x ​ ( z + δ ) ) ≤ ‖ δ ‖ ∞ 2 2 D_{KL}(softmax(z)\parallel softmax(z+\delta))\leq\frac{\parallel\delta\parallel_{\infty}^{2}}{2} 

 Proof. Let P = s ​ o ​ f ​ t ​ m ​ a ​ x ​ ( z ) P=softmax(z) and Q = s ​ o ​ f ​ t ​ m ​ a ​ x ​ ( z + δ ) Q=softmax(z+\delta) .
Write out the KL divergence explicitly:

 D K ​ L ​ ( P ∥ Q ) = ∑ v ∈ 𝒱 P ​ ( v ) ​ l ​ o ​ g ​ P ​ ( v ) Q ​ ( v ) D_{KL}(P\parallel Q)=\sum_{v\mathcal{\in V}}P(v)log\frac{P(v)}{Q(v)} 

Substituting P ​ ( v ) = exp ⁡ ( z v ) Z P(v)=\frac{\exp\left(z_{v}\right)}{Z} and
 Q ​ ( v ) = exp ⁡ ( z v + δ v ) Z ′ Q(v)=\frac{\exp\left(z_{v}+\delta_{v}\right)}{Z^{\prime}} where
 Z = ∑ j exp ⁡ ( z j ) Z=\sum_{j}\exp(z_{j}) and
 Z ′ = ∑ j exp ⁡ ( z j + δ j ) Z^{\prime}=\sum_{j}\exp(z_{j}+\delta_{j}) :

 D K ​ L ​ ( P ∥ Q ) = ∑ v ∈ 𝒱 P ​ ( v ) ​ ( log ⁡ P ​ ( v ) − log ⁡ Q ​ ( v ) ) D_{KL}(P\parallel Q)=\sum_{v\mathcal{\in V}}P(v)\left(\log{P(v)}-\log{Q(v)}\right) 

 D K ​ L ​ ( P ∥ Q ) = ∑ v P ​ ( v ) ​ [ z v − l ​ o ​ g ​ Z − z v − δ v + l ​ o ​ g ​ Z ′ ] = − 𝔼 P ​ [ δ ] + l ​ o ​ g ​ Z ′ − l ​ o ​ g ​ Z D_{KL}(P\parallel Q)=\sum_{v}P(v)\left[z_{v}-logZ-z_{v}-\delta_{v}+logZ^{\prime}\right]=-\mathbb{E}_{P}[\delta]+logZ^{\prime}-logZ 

 Z ′ Z = ∑ j e z j ​ e δ j ∑ j e z j = ∑ j e z j ∑ j e z j ​ e δ j = ∑ j P j ​ e ​ x ​ p ​ ( δ j ) \frac{Z^{\prime}}{Z}=\frac{\sum_{j}{e^{z_{j}}e^{\delta_{j}}}}{\sum_{j}e^{z_{j}}}=\sum_{j}{\frac{e^{z_{j}}}{\sum_{j}e^{z_{j}}}e^{\delta_{j}}}=\sum_{j}P_{j}exp(\delta_{j}) 

Since
 Z ′ Z = ∑ j P ​ ( j ) ​ e ​ x ​ p ​ ( δ j ) = 𝔼 P ​ [ e ​ x ​ p ​ ( δ ) ] \frac{Z^{\prime}}{Z}=\sum_{j}P(j)exp(\delta_{j})=\mathbb{E}_{P}[exp(\delta)] ,
we obtain:

 D K ​ L ​ ( P ∥ Q ) = l ​ o ​ g ​ 𝔼 P ​ [ e ​ x ​ p ​ ( δ ) ] − 𝔼 P ​ [ δ ] D_{KL}(P\parallel Q)=log\mathbb{E}_{P}[exp(\delta)]-\mathbb{E}_{P}[\delta] 

This is the centered cumulant generating function of δ \delta under
 P P .

Define
 ψ ​ ( s ) = l ​ o ​ g ​ 𝔼 P ​ [ e ​ x ​ p ​ ( s ​ δ ) ] − s ⋅ 𝔼 P ​ [ δ ] \psi(s)=log\mathbb{E}_{P}[exp(s\delta)]-s\cdot\mathbb{E}_{P}[\delta] 
for s ∈ [ 0 , 1 ] s\in[0,1] . Then ψ ​ ( 0 ) = 0 \psi(0)=0 and
 ψ ​ ( 1 ) = D K ​ L ​ ( P ∥ Q ) \psi(1)=D_{KL}(P\parallel Q) .

Computing derivatives:

 • 

 ψ ′ ​ ( s ) = 𝔼 Q s ​ [ δ ] − 𝔼 P ​ [ δ ] \psi^{\prime}(s)=\mathbb{E}_{Q_{s}}[\delta]-\mathbb{E}_{P}[\delta] ,
where Q s ​ ( v ) ∝ P ​ ( v ) ​ e ​ x ​ p ​ ( s ⋅ δ v ) Q_{s}(v)\propto P(v)exp(s\cdot\delta_{v}) is the
exponentially tilted distribution.

 • 

 ψ ′′ ​ ( s ) = V ​ a ​ r Q s ​ ( δ ) \psi^{\prime\prime}(s)={Var}_{Q_{s}}(\delta) 

By Taylor’s theorem with integral remainder:

 ψ ​ ( 1 ) = ∫ 0 1 ( 1 − s ) ​ ψ ′′ ​ ( s ) ​ 𝑑 s = ∫ 0 1 ( 1 − s ) ​ V ​ a ​ r Q s ​ ( δ ) ​ 𝑑 s \psi(1)=\int_{0}^{1}(1-s)\,\psi^{\prime\prime}(s)\,ds=\int_{0}^{1}(1-s)\,{Var}_{Q_{s}}(\delta)\,ds 

For any distribution over a bounded domain,
 V ​ a ​ r ​ ( X ) ≤ 𝔼 ​ [ X 2 ] ≤ ‖ X ‖ ∞ 2 Var(X\mathbb{)\leq E[}X^{2}]\leq\parallel X\parallel_{\infty}^{2} .
Since each component δ v \delta_{v} satisfies
 | δ v | ≤ ‖ δ ‖ ∞ |\delta_{v}|\leq\parallel\delta\parallel_{\infty} :

 D K ​ L ​ ( P ∥ Q ) ≤ ‖ δ ‖ ∞ 2 ​ ∫ 0 1 ( 1 − s ) ​ 𝑑 s = ‖ δ ‖ ∞ 2 2 ■ D_{KL}(P\parallel Q)\leq\parallel\delta\parallel_{\infty}^{2}\int_{0}^{1}(1-s)\,ds=\frac{\parallel\delta\parallel_{\infty}^{2}}{2}\quad\quad\blacksquare 

 A.3 Assumption 1 (Local Lipschitz
Continuity)

On the domain of bounded inputs (finite embeddings, finite sequence
length), the logit function f θ f_{\theta} is locally Lipschitz
continuous: there exists L θ < ∞ L_{\theta}<\infty such that

 ‖ f θ ​ ( c 1 ) − f θ ​ ( c 2 ) ‖ ∞ ≤ L θ ⋅ d ​ ( c 1 , c 2 ) \parallel f_{\theta}(c_{1})-f_{\theta}(c_{2})\parallel_{\infty}\leq L_{\theta}\cdot d(c_{1},c_{2}) 

for all valid contexts c 1 , c 2 c_{1},c_{2} in the bounded domain, where
 d ​ ( ⋅ , ⋅ ) d(\cdot,\cdot) is a distance metric on the input space.

 Justification. This is a standard property for neural networks
on compact domains. For transformers specifically, Kim et al. ( 2021 ) 
establish local Lipschitz bounds for self-attention on bounded input
domains. Recent work shows the per-layer local Lipschitz constant scales
as O ​ ( C ​ n ) O(C\sqrt{n}) where n n is sequence length and C C depends on
weight norms and attention distribution concentration.

 A.4 Proposition 1 (Realizability Gap
Comparison)

 A.4.1 Part I – Same-Model
Bound

 Statement. Under Assumption 1, the per-position
realizability gap for same-model privileged distillation satisfies: 

 D K ​ L ​ ( P T ∥ P S ) ≤ L θ 2 ⋅ Δ ​ ( g ) 2 2 D_{KL}(P_{T}\parallel P_{S})\leq\frac{L_{\theta}^{2}\cdot\Delta(g)^{2}}{2} 

 where Δ ​ ( g ) = d ​ ( c T , c S ) \Delta(g)=d(c_{T},c_{S}) is the input-space
distance attributable to the ground truth tokens g g . 

 Proof. Teacher and student evaluate the same function
 f θ f_{\theta} on different inputs c T c_{T} and c S c_{S} . By
Assumption 1:

 ‖ z T − z S ‖ ∞ = ‖ f θ ​ ( c T ) − f θ ​ ( c S ) ‖ ∞ ≤ L θ ⋅ d ​ ( c T , c S ) = L θ ⋅ Δ ​ ( g ) \parallel z_{T}-z_{S}\parallel_{\infty}=\parallel f_{\theta}(c_{T})-f_{\theta}(c_{S})\parallel_{\infty}\leq L_{\theta}\cdot d(c_{T},c_{S})=L_{\theta}\cdot\Delta(g) 

Applying Lemma 1 with δ = z T − z S \delta=z_{T}-z_{S} :

 D K ​ L ​ ( P T ∥ P S ) ≤ ‖ z T − z S ‖ ∞ 2 2 ≤ L θ 2 ⋅ Δ ​ ( g ) 2 2 ■ D_{KL}(P_{T}\parallel P_{S})\leq\frac{\parallel z_{T}-z_{S}\parallel_{\infty}^{2}}{2}\leq\frac{L_{\theta}^{2}\cdot\Delta(g)^{2}}{2}\quad\quad\blacksquare 

 Key properties of this bound: 

 1. 

It depends only on the model θ \theta (via L θ L_{\theta} ) and the
information content of the ground truth g g (via Δ ​ ( g ) \Delta(g) ).

 2. 

It does not depend on any capacity gap between two different models.

 3. 

The difficulty of distillation scales with how much the ground truth
changes the prediction, not with architectural differences.

 A.4.2 Part II – Cross-Model
Comparison

 Statement. Under Assumption 1, for cross-model
distillation with teacher ϕ \phi and student θ \theta ,
where the teacher receives privileged input c T c_{T} and the
student receives c S c_{S} , the realizability gap satisfies: 

 D K ​ L ( P ϕ ( ⋅ | c T ) ∥ P θ ( ⋅ | c S ) ) ≤ ( L θ ⋅ Δ ​ ( g ) + ‖ f ϕ ​ ( c T ) − f θ ​ ( c T ) ‖ ∞ ) 2 2 D_{KL}(P_{\phi}(\cdot|c_{T})\parallel P_{\theta}(\cdot|c_{S}))\leq\frac{\left(L_{\theta}\cdot\Delta(g)+\left\|f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{T}\right)\right\|_{\infty}\right)^{2}}{2} 

 Proof. The logit difference between the cross-model teacher and
the student decomposes as:

 f ϕ ​ ( c T ) − f θ ​ ( c S ) = [ f ϕ ​ ( c T ) − f θ ​ ( c T ) ] + [ f θ ​ ( c T ) − f θ ​ ( c S ) ] f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{S}\right)=\left[f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{T}\right)\right]+\left[f_{\theta}\left(c_{T}\right)-f_{\theta}\left(c_{S}\right)\right] 

The first bracket captures the model mismatch: the difference in logit
functions between models ϕ \phi and θ \theta evaluated on the same
input. The second bracket is the same input perturbation that appears in
Part I. By the triangle inequality:

 ‖ f ϕ ​ ( c T ) − f θ ​ ( c S ) ‖ ∞ ≤ ‖ f ϕ ​ ( c T ) − f θ ​ ( c T ) ‖ ∞ + L θ ⋅ Δ ​ ( g ) \left\|f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{S}\right)\right\|_{\infty}\leq\left\|f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{T}\right)\right\|_{\infty}+L_{\theta}\cdot\Delta(g) 

where the second inequality applies Assumption 1 to the student model
 θ \theta . Applying Lemma 1 with
 δ = f ϕ ​ ( c T ) − f θ ​ ( c S ) \delta=f_{\phi}\left(c_{T}\right)-f_{\theta}\left(c_{S}\right) 
yields the bound. ■ \blacksquare 

 A.5 Remark

The bound in Part I may be quantitatively loose: the Lipschitz constant
 L θ L_{\theta} for a deep transformer can be large due to multiplicative
composition across layers. However, the comparison in Part II does not
depend on the tightness of the Lipschitz bound. The same
 L θ ⋅ Δ ​ ( g ) L_{\theta}\cdot\Delta(g) term appears in both bounds, and the
cross-model bound is strictly larger by the model-mismatch term
regardless of the absolute magnitudes. A loose Lipschitz constant
affects both bounds equally — it does not favor one distillation
setting over the other.

Moreover, L θ L_{\theta} is the same for both teacher and student in the
same-model case (since they share parameters), so it does not introduce
a capacity mismatch — it merely scales the bound uniformly.

 Appendix B Proposition 2 Full
Proof

 Proposition 2 ( R = 1 R=1 Filtering Yields the RL-Optimal
Policy). We provide the full derivation here. For binary reward:

 For binary reward R ​ ( τ ) ∈ { 0 , 1 } R(\tau)\ \in\ \{0,\ 1\} , the
exponential factor takes only two values: 

 exp ⁡ ( R ​ ( τ ) β ) = e ​ x ​ p ​ ( 1 β ) i ​ f ​ R ​ ( τ ) = 1 , a ​ n ​ d e ​ x ​ p ​ ( R ​ ( τ ) β ) = 1 i ​ f ​ R ​ ( τ ) = 0 \exp\left(\frac{R(\tau)}{\beta}\right)\ =\ exp\left(\frac{1}{\beta}\right)\ \ if\ R(\tau)\ =\ 1,\ \ \ \ \ and\ \ \ \ \ exp\left(\frac{R(\tau)}{\beta}\right)\ =\ 1\ \ if\ R(\tau)\ =\ 0 

 Substituting into the Gibbs distribution: 

 π ∗ ​ ( τ ) = π r ​ e ​ f ​ ( τ ) ⋅ e ​ x ​ p ​ ( 1 ​ [ R ​ ( τ ) = 1 ] β ) Z ​ ( β ) \pi^{*}(\tau)\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ exp\left(\frac{1\left[R(\tau)=\ 1\right]}{\beta}\right)}{Z(\beta)} 

 where the partition function evaluates to: 

 Z ​ ( β ) = P π r ​ e ​ f ​ ( R = 0 ) ⋅ 1 + P π r ​ e ​ f ​ ( R = 1 ) ⋅ e ​ x ​ p ​ ( 1 β ) = ( 1 − p ) + p ⋅ e ​ x ​ p ​ ( 1 β ) Z(\beta)\ =\ P_{\pi_{ref}}(R\ =\ 0)\ \cdot\ 1\ +\ P_{\pi_{ref}}(R\ =\ 1)\ \cdot\ exp\left(\frac{1}{\beta}\right)\ =\ (1\ -\ p)\ +\ p\ \cdot\ exp\left(\frac{1}{\beta}\right) 

 where p = P π ​ _ ​ r ​ e ​ f ​ ( R = 1 ) p\ =\ P_{\pi\_ref}(R\ =\ 1) is the
probability that π r ​ e ​ f \pi_{ref} generates a correct trajectory.
This gives the relative weighting of correct vs incorrect trajectories: 

 π ∗ ​ ( τ | R = 1 ) π ∗ ​ ( τ | R = 0 ) = e ​ x ​ p ​ ( 1 β ) ( for trajectories with equal  ​ π ref ​  weight ) \frac{\pi^{*}(\tau\ |\ R\ =\ 1)}{\pi^{*}(\tau\ |\ R\ =\ 0)}\ =\ exp\left(\frac{1}{\beta}\right)\ \ \ \ \ (\text{for trajectories with equal }\pi_{\text{ref}}\text{ weight}) 

 The hard-threshold limit. As 
 β → 0 + \beta\ \rightarrow\ 0^{+} , the exponential ratio 
 e ​ x ​ p ​ ( 1 / β ) → ∞ exp(1/\beta)\ \rightarrow\ \infty . This means correct
trajectories receive infinitely more weight than incorrect ones.
Concretely, for any incorrect trajectory with 
 R ​ ( τ ) = 0 R(\tau)\ =\ 0 : 

 π ∗ ​ ( τ ) = π r ​ e ​ f ​ ( τ ) ⋅ e ​ x ​ p ​ ( 1 ​ [ 0 = 1 ] β ) Z ​ ( β ) = π r ​ e ​ f ​ ( τ ) ⋅ 1 Z ​ ( β ) → 0 a ​ s ​ β → 0 + \pi^{*}(\tau)\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ exp\left(\frac{1[0=\ 1]}{\beta}\right)}{Z(\beta)}\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ 1}{Z(\beta)}\ \rightarrow\ 0\ \ \ \ \ as\ \beta\ \rightarrow\ 0^{+} 

 because Z ​ ( β ) ∼ p ⋅ e ​ x ​ p ​ ( 1 / β ) Z(\beta)\ \sim\ p\ \cdot\ exp(1/\beta) 
 grows without bound while the numerator is fixed. Conversely, for
a correct trajectory with R ​ ( τ ) = 1 R(\tau)\ =\ 1 : 

 π ∗ ​ ( τ ) = π r ​ e ​ f ​ ( τ ) ⋅ e ​ x ​ p ​ ( 1 ​ [ 1 = 1 ] β ) Z ​ ( β ) = π r ​ e ​ f ​ ( τ ) ⋅ e ​ x ​ p ​ ( 1 β ) p ⋅ e ​ x ​ p ​ ( 1 β ) + ( 1 − p ) → π r ​ e ​ f ​ ( τ ) p = π r ​ e ​ f ​ ( τ ) P π ​ _ ​ r ​ e ​ f ​ ( R = 1 ) \pi^{*}(\tau)\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ exp\left(\frac{1[1\ =\ 1]}{\beta}\right)}{Z(\beta)}\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ exp\left(\frac{1}{\beta}\right)}{p\ \cdot\ exp\left(\frac{1}{\beta}\right)\ +\ (1\ -\ p)}\ \rightarrow\ \frac{\pi_{ref}(\tau)}{p}\ =\ \frac{\pi_{ref}(\tau)}{P_{\pi\_ref}(R\ =\ 1)} 

 where the limit follows from dividing numerator and denominator
by e ​ x ​ p ​ ( 1 / β ) exp(1/\beta) and noting that 
 ( 1 − p ) / e ​ x ​ p ​ ( 1 / β ) → 0 (1\ -\ p)\ /\ exp(1/\beta)\ \rightarrow\ 0 . Therefore, in
the limit: 

 π ∗ ​ ( τ ) → π r ​ e ​ f ​ ( τ | R ​ ( τ ) = 1 ) = π r ​ e ​ f ​ ( τ ) ⋅ 1 ​ [ R ​ ( τ ) = 1 ] P π ​ _ ​ r ​ e ​ f ​ ( R = 1 ) \pi^{*}(\tau)\ \rightarrow\ \pi_{ref}(\tau\ |\ R(\tau)\ =\ 1)\ =\ \frac{\pi_{ref}(\tau)\ \cdot\ 1[R(\tau)\ =\ 1]}{P_{\pi\_ref}(R\ =\ 1)} 

 Appendix C Experimental
Details

 Table 2: Hyperparameters used in all experiments. 

 Category 
 Parameter 
 Value 

 Model 
 Architecture 
 Qwen2.5-Math-1.5B-Instruct 

 Precision 
 bfloat16 

 Max sequence length 
 4096 

 Dataset 
 Training 
 OpenMathInstruct-2 

 Validation samples 
 2048 

 Reward function 
 hf_math_verify (binary) 

 Optimization 
 Optimizer 
 AdamW ( β 1 = 0.9 \beta_{1}{=}0.9 , β 2 = 0.999 \beta_{2}{=}0.999 , ϵ = 10 − 8 \epsilon{=}10^{-8} ) 

 Learning rate 
 10 − 6 10^{-6} 

 LR schedule 
 Linear warmup (50 steps, 0.1 → \to 1.0 × \times ), then constant 

 Weight decay 
 0.01 

 Max gradient norm 
 1.0 

 GRPO 
 Prompts per step 
 32 

 Generations per prompt 
 16 

 Training batch size 
 512 

 Advantage normalization 
 Leave-one-out baseline 

 Ratio clip ϵ \epsilon 

 0.2 

 Loss granularity 
 Token-level 

 Reference KL penalty β \beta 

 0.0 

 Generation temperature 
 1.0 

 Generation backend 
 vLLM (colocated) 

 Sequence packing 
 Enabled 

 Privileged Distillation 
 Distillation weight λ \lambda 

 {0.01, 0.1} 

 Distillation loss 
 JSD 

 Teacher type 
 {Drifting (shared current policy weights), Frozen (initial weights)} 

 Normalization 
 Global token count (rank-invariant) 

 Teacher top- k k 

 64 

 Cliff threshold 
 reward_sum = 0.0 =0.0 

 Teacher success threshold 
 reward ≥ 1.0 \geq 1.0 

 Max cliff prompts per step 
 32 

 Evaluation 
 Validation period 
 Every 10 steps 

 Reported metrics 
 pass@1, pass@4, pass@8 

 Infrastructure 
 GPUs 
 8 × \times H200 (1 node); replicated on 8 × \times H100 

 Seed 
 42 

 Appendix D Hardware Variation

All main results (Table  1 ) were obtained on 8 × \times H200 GPUs. We replicated all five configurations on 8 × \times H100 GPUs with identical hyperparameters. The training computation is mathematically equivalent; differences arise from floating-point non-determinism across GPU microarchitectures (different matmul tiling, reduction order, and FlashAttention kernel implementations), which compound over 2000 training steps to produce subtly different final weights.

 Table 3: Best pass@ k k on 8 × \times H100 GPUs. Same setup as Table  1 .
Bold indicates best in column. 

 Method 
 pass@1 
 pass@4 
 pass@8 

 GRPO Baseline 
 0.6509 
 0.7739 
 0.8223 

 HDPO (frozen, λ \lambda =0.01) 
 0.6484 
 0.7773 
 0.8252 

 HDPO (frozen, λ \lambda =0.1) 
 0.6343 
 0.7856 
 0.8369 

 HDPO (drifting, λ \lambda =0.01) 
 0.6499 
 0.7783 
 0.8213 

 HDPO (drifting, λ \lambda =0.1) 
 0.6343 
 0.7832 
 0.8359 

The qualitative findings are consistent: λ = 0.1 \lambda{=}0.1 improves pass@8 by + 1.4 +1.4 – 1.5 % 1.5\% over baseline on both hardware configurations, and pass@1 remains highest for the baseline and λ = 0.01 \lambda{=}0.01 configurations. The primary quantitative difference is at λ = 0.01 \lambda{=}0.01 : on H200, drifting- 0.01 0.01 achieves the highest pass@4 ( + 1.1 % +1.1\% ), while on H100 the improvement is smaller ( + 0.4 % +0.4\% ) and the best pass@4 shifts to frozen- 0.1 0.1 . This suggests the λ = 0.01 \lambda{=}0.01 improvements, while directionally consistent, are close to the noise floor introduced by hardware-level floating-point non-determinism and temperature-1 evaluation variance. The λ = 0.1 \lambda{=}0.1 coverage improvements are robust across both configurations.

LLM Usage Statement

The research idea underlying HDPO—using privileged self-distillation to provide
learning signal on cliff prompts—originated with the author. However, an AI
language model (Claude, Anthropic) was used extensively throughout this project in
ways that go beyond minor writing assistance. Specifically: (1) the mathematical
formalization of Proposition 1 and Proposition 2, including the proof structure and
verification of correctness, was developed collaboratively with LLM assistance;
(2) the paper text, including the related work survey, theoretical exposition, and
discussion sections, was substantially drafted and edited with LLM assistance;
and (3) the LLM was used as a research collaborator to brainstorm explanations
for experimental phenomena (e.g., the pass@1 vs. pass@ k k tradeoff) and to
compare HDPO against related work.
All experimental results (training runs, metric measurements) were produced
by the author without LLM involvement. The LLM also provided substantial
code implementation assistance for the training infrastructure.

References

 Agarwal et al. (2024) 

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos,
Matthieu Geist, and Olivier Bachem.

 On-policy distillation of language models: Learning from
self-generated mistakes.

 In International Conference on Learning Representations , 2024.

 URL https://arxiv.org/abs/2306.13649 .

 Cui et al. (2025) 

Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, et al.

 Process reinforcement through implicit rewards.

 arXiv preprint arXiv:2502.01456 , 2025.

 URL https://arxiv.org/abs/2502.01456 .

 DeepSeek-AI et al. (2025) 

DeepSeek-AI, Daya Guo, Dejian Yang, et al.

 DeepSeek-R1: Incentivizing reasoning capability in LLMs via
reinforcement learning.

 arXiv preprint arXiv:2501.12948 , 2025.

 URL https://arxiv.org/abs/2501.12948 .

 Dou et al. (2025) 

Shihan Dou, Muling Wu, Jingwen Xu, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing
Huang.

 Improving RL exploration for LLM reasoning through retrospective
replay.

 arXiv preprint arXiv:2504.14363 , 2025.

 URL https://arxiv.org/abs/2504.14363 .

 Hendrycks et al. (2021) 

Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric
Tang, Dawn Song, and Jacob Steinhardt.

 Measuring mathematical problem solving with the MATH dataset.

 In Neural Information Processing Systems , 2021.

 URL https://arxiv.org/abs/2103.03874 .

 Hinton et al. (2015) 

Geoffrey Hinton, Oriol Vinyals, and Jeff Dean.

 Distilling the knowledge in a neural network.

 arXiv preprint arXiv:1503.02531 , 2015.

 URL https://arxiv.org/abs/1503.02531 .

 Hübotter et al. (2026) 

Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco
Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening,
Carlos Guestrin, and Andreas Krause.

 Reinforcement learning via self-distillation.

 arXiv preprint arXiv:2601.20802 , 2026.

 URL https://arxiv.org/abs/2601.20802 .

 Jiang et al. (2025) 

Guochao Jiang, Wenfeng Feng, Guofeng Quan, Chuzhan Hao, Yuewei Zhang, Guohua
Liu, and Hao Wang.

 VCRL: Variance-based curriculum reinforcement learning for large
language models.

 arXiv preprint arXiv:2509.19803 , 2025.

 URL https://arxiv.org/abs/2509.19803 .

 Kim et al. (2021) 

Hyunjik Kim, George Papamakarios, and Andriy Mnih.

 The Lipschitz constant of self-attention.

 In International Conference on Machine Learning , 2021.

 URL https://arxiv.org/abs/2006.04710 .

 Kwon et al. (2023) 

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu,
Joseph E. Gonzalez, Hao Zhang, and Ion Stoica.

 Efficient memory management for large language model serving with
PagedAttention.

 In Symposium on Operating Systems Principles , 2023.

 URL https://arxiv.org/abs/2309.06180 .

 Le et al. (2025) 

Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Lai, and Eunho Yang.

 No prompt left behind: Exploiting zero-variance prompts in LLM
reinforcement learning via entropy-guided advantage shaping.

 arXiv preprint arXiv:2509.21880 , 2025.

 URL https://arxiv.org/abs/2509.21880 .

 Li et al. (2025) 

Long Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu, Yanting Miao, Wei Pang,
Xiaoyu Tan, Wei Chu, Zhe Wang, Shirui Pan, Chao Qu, and Yuan Qi.

 The choice of divergence: A neglected key to mitigating diversity
collapse in reinforcement learning with verifiable reward.

 arXiv preprint arXiv:2509.07430 , 2025.

 URL https://arxiv.org/abs/2509.07430 .

 Lightman et al. (2024) 

Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy
Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.

 Let’s verify step by step.

 In International Conference on Learning Representations , 2024.

 URL https://arxiv.org/abs/2305.20050 .

 Lin (1991) 

Jianhua Lin.

 Divergence measures based on the Shannon entropy.

 IEEE Transactions on Information Theory , 37(1):145–151, 1991.

 URL https://ieeexplore.ieee.org/document/61115/ .

 Liu et al. (2025) 

Huanyu Liu, Jia Li, Chang Yu, Taozhi Chen, Yihong Dong, Lecheng Wang, Xiaolong
Hu, and Ge Li.

 EvoCoT: Overcoming the exploration bottleneck in reinforcement
learning.

 arXiv preprint arXiv:2508.07809 , 2025.

 URL https://arxiv.org/abs/2508.07809 .

 Loshchilov and Hutter (2019) 

Ilya Loshchilov and Frank Hutter.

 Decoupled weight decay regularization.

 In International Conference on Learning Representations , 2019.

 URL https://arxiv.org/abs/1711.05101 .

 Ma et al. (2025) 

Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo
Niu, Chengyu Shen, Runming He, Yanhao Li, Bin Cui, and Wentao Zhang.

 Learning what reinforcement learning can’t: Interleaved online
fine-tuning for hardest questions.

 arXiv preprint arXiv:2506.07527 , 2025.

 URL https://arxiv.org/abs/2506.07527 .

 Ouyang et al. (2022) 

Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela
Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.

 Training language models to follow instructions with human feedback.

 In Neural Information Processing Systems , 2022.

 URL https://arxiv.org/abs/2203.02155 .

 Penaloza et al. (2026) 

Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste,
Laurent Charlin, and Massimo Caccia.

 Privileged information distillation for language models.

 arXiv preprint arXiv:2602.04942 , 2026.

 URL https://arxiv.org/abs/2602.04942 .

 Rafailov et al. (2023) 

Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D.
Manning, and Chelsea Finn.

 Direct preference optimization: Your language model is secretly a
reward model.

 In Neural Information Processing Systems , 2023.

 URL https://arxiv.org/abs/2305.18290 .

 Schulman et al. (2017) 

John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov.

 Proximal policy optimization algorithms.

 arXiv preprint arXiv:1707.06347 , 2017.

 URL https://arxiv.org/abs/1707.06347 .

 Shao et al. (2024) 

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei
Zhang, Mingchuan Zhang, Y.K. Li, Y. Wu, and Daya Guo.

 DeepSeekMath: Pushing the limits of mathematical reasoning in open
language models.

 arXiv preprint arXiv:2402.03300 , 2024.

 URL https://arxiv.org/abs/2402.03300 .

 Toshniwal et al. (2024) 

Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan
Ayrapetyan, and Igor Gitman.

 OpenMathInstruct-2: Accelerating AI for math with massive
open-source instruction data.

 arXiv preprint arXiv:2410.01560 , 2024.

 URL https://arxiv.org/abs/2410.01560 .

 Vapnik and Vashist (2009) 

Vladimir Vapnik and Akshay Vashist.

 A new learning paradigm: Learning using privileged information.

 Neural Networks , 22(5–6):544–557, 2009.

 doi: 10.1016/j.neunet.2009.06.042 .

 Wang et al. (2025) 

Xinyi Wang, Jinyi Han, Zishang Jiang, Tingyun Li, Jiaqing Liang, Sihang Jiang,
Zhaoqian Dai, Shuguang Ma, Fei Yu, and Yanghua Xiao.

 HINT: Helping ineffective rollouts navigate towards effectiveness.

 arXiv preprint arXiv:2510.09388 , 2025.

 URL https://arxiv.org/abs/2510.09388 .

 Xu et al. (2025) 

Hongling Xu, Qi Zhu, Heyuan Deng, Jinpeng Li, Lu Hou, Yasheng Wang, Lifeng
Shang, Ruifeng Xu, and Fei Mi.

 KDRL: Post-training reasoning LLMs via unified knowledge
distillation and reinforcement learning.

 arXiv preprint arXiv:2506.02208 , 2025.

 URL https://arxiv.org/abs/2506.02208 .

 Yang et al. (2024) 

An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li,
Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng
Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang.

 Qwen2.5-Math technical report: Toward mathematical expert model via
self-improvement.

 arXiv preprint arXiv:2409.12122 , 2024.

 URL https://arxiv.org/abs/2409.12122 .

 Yang et al. (2026) 

Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin.

 Learning beyond teacher: Generalized on-policy distillation with
reward extrapolation.

 arXiv preprint arXiv:2602.12125 , 2026.

 URL https://arxiv.org/abs/2602.12125 .

 Yu et al. (2025) 

Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al.

 DAPO: An open-source LLM reinforcement learning system at scale.

 arXiv preprint arXiv:2503.14476 , 2025.

 URL https://arxiv.org/abs/2503.14476 .

 Zhang et al. (2025a) 

Hongzhi Zhang, Jia Fu, Jingyuan Zhang, Kai Fu, Qi Wang, Fuzheng Zhang, and
Guorui Zhou.

 RLEP: Reinforcement learning with experience replay for LLM
reasoning.

 arXiv preprint arXiv:2507.07451 , 2025a.

 URL https://arxiv.org/abs/2507.07451 .

 Zhang et al. (2025b) 

Xichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan, Shaozuo Yu, Ziyi He, and Jiaya
Jia.

 Scaf-GRPO: Scaffolded group relative policy optimization for
enhancing LLM reasoning.

 arXiv preprint arXiv:2510.19807 , 2025b.

 URL https://arxiv.org/abs/2510.19807 .

 Zhang et al. (2026) 

Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo
Yang, Zhuowen Tu, Wei Xia, and Stefano Soatto.

 Reinforcement-aware knowledge distillation for LLM reasoning.

 arXiv preprint arXiv:2602.22495 , 2026.

 URL https://arxiv.org/abs/2602.22495 .

 Zhao et al. (2026) 

Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and
Aditya Grover.

 Self-distilled reasoner: On-policy self-distillation for large
language models.

 arXiv preprint arXiv:2601.18734 , 2026.

 URL https://arxiv.org/abs/2601.18734 .

 Experimental support, please
 view the build logs 
 for errors. Generated by

 L
 A 
 T
 E 

 xml 

 .

Instructions for reporting errors

We are continuing to improve HTML versions of papers, and your feedback helps enhance accessibility and mobile
 support. To report errors in the HTML that will help us improve conversion and rendering, choose any of the
 methods listed below:

Click the "Report Issue" ( 

 ) button, located in the page header.

 Tip: You can select the relevant text first, to include it in your report.

Our team has already identified the following issues . We appreciate your time reviewing and reporting rendering errors we
 may not have found yet. Your efforts will help us improve the HTML versions for all readers, because disability
 should not be a barrier to accessing research. Thank you for your continued support in championing open access for
 all.

Have a free development cycle? Help support accessibility at arXiv! Our collaborators at LaTeXML maintain a list of packages that need conversion , and welcome developer contributions .

BETA

