Research notes · September 26, 2026

Online
InterventionOLIVE

Online intervention improves how students learn from stronger teachers. Let the student try, have the teacher continue from there, and train on what the teacher does next.

Project main contributors

Haojin Wang Dylan Zhang (Project Lead)

With collaborators at MIT · MIT-IBM Watson Lab · University of Pennsylvania (UPenn) · University of Washington (UW)

Help at the point where learning gets hard.

A teacher can solve a problem beautifully and still give a student a difficult example to learn from. The student makes different choices, ends up in different situations, and needs help with what follows its own attempt. Training on a fixed collection of complete teacher solutions can leave those situations uncovered.

On-policy distillation (OPD) takes this issue seriously: sample from the student, then use the teacher to supervise those samples. Generalized Knowledge Distillation is one example of this approach. We like the principle of meeting the student where it is.

Many OPD approaches may be unstable when the teacher and student are poorly matched. Recent studies trace failures to mismatch between the teacher and student and some approximations to the token-level objective. Online intervention offers a direct way to supply a connected teacher continuation from the student’s own context. See Zhu et al. and Li et al.

There is also a trajectory problem. In token-level OPD, the teacher can recommend a correction at one position while later targets are still conditioned on the student’s original continuation. The student receives advice without seeing the path that follows that advice. Trajectory-Refined Distillation studies this related problem of prefix failure.

We explored a direct way to supply that path: let the teacher take over from the student’s current context and generate a continuation. We call this Online Intervention, or OLIVE.

The teacher continues. The student learns.

OLIVE repeats a short loop throughout training. The current student begins an attempt. A teacher continues from that prefix. We update the student with ordinary cross-entropy loss on the teacher’s continuation, then collect fresh prefixes from the updated student.

Student prefix

Context the student actually reaches

Teacher continuation

A path forward from that context

Masked loss

Learn only from teacher outputs

Updated student

New attempts, new supervision

Prompt + student prefix · context onlyTeacher continuation · learning signal

The student starts with its current policy. Prefixes are collected during training, so they reflect the behavior we are trying to improve.

The teacher receives the prompt and the student’s attempt. Each new teacher token conditions on its own previous tokens, creating a connected continuation.

The student sees the full sequence. The prompt and student prefix are masked out of the loss. Cross-entropy teaches the student to produce the teacher’s continuation in that context.

After updating the student, we collect new attempts and ask for new teacher continuations. The situations where we ask for help evolve with the learner.

For an agent, the prefix is a history of actions and observations. The teacher takes over interaction with the environment for a few turns, and its actions produce the next observations. We keep that history as context and train only on the teacher’s actions. Student turns and environment observations receive no loss.

The teacher only needs to return text. Its output can be retokenized for the student, so this procedure does not require teacher logits or a shared tokenizer. We can also use a short continuation without asking the teacher to finish or verify the whole solution.

We currently hand control to the teacher after a fixed prefix length. The teacher then demonstrates a continuation from that state. Recovery depends on the prefix still admitting a useful continuation.

This draws on a long history of learning from expert help at learner-visited states, including DAgger and On-Policy Expert Corrections. Our focus here is on refreshing those intervention contexts throughout training and learning from the resulting text.

A useful gain on hard reasoning.

We tested reasoning on 18 synthetic RLVE environments, with difficulty chosen to challenge the students. Both students use thinking, and the teacher is Qwen3-4B-Thinking-2507. OLIVE achieves the highest average accuracy among these methods for both student sizes.

Reasoning · 18 RLVE environments

Higher average accuracy for both students

Teacher: Qwen3-4B-Thinking-2507 · Higher is better.

Avg@8

Shared scale · 0–30%

Qwen3-1.7B

  1. Original student3.3%
  2. Offline teacher data5.6%
  3. OPD4.6%
  4. OLIVE7.8%

Avg@8 (%)

Qwen3-4B

  1. Original student18.1%
  2. Offline teacher data19.9%
  3. OPD21.3%
  4. OLIVE23.5%

Avg@8 (%)

Pass@8

Shared scale · 0–60%

Qwen3-1.7B

  1. Original student11.1%
  2. Offline teacher data15.0%
  3. OPD14.4%
  4. OLIVE19.4%

Pass@8 (%)

Qwen3-4B

  1. Original student46.1%
  2. Offline teacher data53.3%
  3. OPD45.6%
  4. OLIVE52.2%

Pass@8 (%)

Avg@8 is average correctness across eight attempts. Pass@8 measures whether at least one of eight attempts succeeds. For each metric, both students share a scale: 0–30% for Avg@8 and 0–60% for Pass@8.

OLIVE uses asynchronous training, a 4,096-token student prefix, and up to 1,024 teacher tokens, without filtering. Baselines use a cap of 7,168 distilled tokens per rollout.

The reasoning OPD baseline uses a top-16 approximation to the teacher distribution.

OLIVE more than doubles the smaller student’s average accuracy. The larger student improves on both metrics relative to its starting point. Across both students, OLIVE achieves the highest avg@8 in the comparison. Offline teacher data has the highest pass@8 for the larger student.

Teacher takeover also helped agents.

Across five interactive environments, OLIVE improves success over both the original student and plain OPD. It also leads the compared student-training methods on four of the five tasks, with Guided OPD slightly ahead on BabyAI. The student is Qwen3-1.7B, learning from Qwen3-32B. These results extend the benefit of teacher continuations to tasks where each action changes what the agent sees next.

Agents · Five interactive environments

Ahead of plain OPD on all five tasks.

Student: Qwen3-1.7B · Teacher: Qwen3-32B

ALFWorld

Success · Avg@4 · 0–100%
  1. Original student19.38%
  2. Teacher reference52.12%
  3. OPD22.25%
  4. TCoD-B2F37.10%
  5. TCoD-F2B28.00%
  6. Guided OPD27.12%
  7. OLIVE40.00%

Success rate (%) · higher is better

ScienceWorld

Success · Avg@4 · 0–100%
  1. Original student0.12%
  2. Teacher reference15.62%
  3. OPD0.00%
  4. TCoD-B2F0.75%
  5. TCoD-F2B0.50%
  6. Guided OPD0.75%
  7. OLIVE7.50%

Success rate (%) · higher is better

TextCraft

Success · Avg@4 · 0–100%
  1. Original student23.00%
  2. Teacher reference85.50%
  3. OPD29.50%
  4. TCoD-B2F39.50%
  5. TCoD-F2B45.50%
  6. Guided OPD45.50%
  7. OLIVE55.25%

Success rate (%) · higher is better

BabyAI

Success · Avg@4 · 0–100%
  1. Original student38.33%
  2. Teacher reference83.33%
  3. OPD43.06%
  4. TCoD-B2F66.30%
  5. TCoD-F2B62.50%
  6. Guided OPD67.78%
  7. OLIVE67.50%

Success rate (%) · higher is better

SearchQA

Success · Avg@4 · 0–100%
  1. Original student30.50%
  2. Teacher reference55.06%
  3. OPD29.56%
  4. TCoD-B2F37.69%
  5. TCoD-F2B37.75%
  6. Guided OPD37.50%
  7. OLIVE39.06%

Success rate (%) · higher is better

Task success rates, averaged over four evaluations (avg@4). Every environment uses the same 0–100% scale; values are printed exactly, including near-zero results.

The hatched teacher reference row reports Qwen3-32B. All other rows report Qwen3-1.7B.

ScienceWorld is the clearest example of learning from a very weak starting point: OLIVE reaches 7.5% success where the original student and OPD are near zero. Teacher takeover provides a productive learning signal even when the student rarely completes the task on its own.

Student rows report the student’s task performance. The teacher is shown separately for reference.

Less waiting for the teacher.

Online supervision adds a practical cost: a teacher has to generate while training runs. We overlap teacher continuation for one batch with student prefix generation for the next. This reduces idle time, at the price of training on prefixes from a slightly older student.

In our reasoning experiment, asynchronous OLIVE reduced training time by 23.8% relative to synchronous OLIVE. The asynchronous depth was three updates. Pass@8 changed from 54.4% to 52.2%, while avg@8 remained similar.

A short continuation and overlapping generation made the approach practical in this setting. This gives us a practical starting point for larger training runs, where balancing throughput and prefix freshness will be especially interesting.

Analysis

Learn more while preserving what the student knows.

Refreshing the student’s own prefixes improves transfer and preserves more of its general capabilities. We compared complete teacher solutions, teacher continuations from a fixed initial student, and continuations collected from the changing student throughout training.

For this analysis, all three training conditions use complete responses filtered by a verifier. The fixed-prefix baseline collects the student’s attempts once. OLIVE refreshes those attempts as the student learns, so each new intervention responds to its current behavior.

01 · Learning on RLVE

Fresh contexts improve learning.

Online prefixes versus fixed data.

Avg@k

0–15% · higher is better
  1. Original student3.3%
  2. Offline full SFT8.1%
  3. Fixed student prefix (OEC)10.7%
  4. OLIVE12.2%

Avg@k (%)

Pass@k

0–40% · higher is better
  1. Original student11.1%
  2. Offline full SFT23.9%
  3. Fixed student prefix (OEC)30.0%
  4. OLIVE33.3%

Pass@k (%)

Avg@k averages correctness over k attempts. Pass@k measures whether at least one succeeds. We retain the source figure’s notation.

This analysis uses complete, verifier-filtered continuations. OEC collects prefixes from a fixed student; OLIVE refreshes prefixes as the student learns.

02 · General capabilities

Smaller drops beyond the training tasks.

Change from the original student.

Zero is the starting point. A shorter bar means a smaller drop in average accuracy.

General benchmark average

Percentage-point change
  1. Original student0.0 pp
  2. Offline full SFT−3.4 pp
  3. Fixed student prefix (OEC)−1.3 pp
  4. OLIVE−0.9 pp

Change in accuracy (percentage points)

Average accuracy change on AIME25, LiveCodeBench v6, IFEval, and GPQA Diamond, relative to the original student. Zero means no change; negative values indicate a drop.

OLIVE’s average drop is 0.9 percentage points, compared with 3.4 for offline full SFT and 1.3 for fixed student prefixes.

OLIVE achieves the highest task performance on both metrics. Collecting teacher continuations from a fixed student already helps. Refreshing those prefixes online adds a further gain, connecting the teacher’s demonstrations to the student that is learning from them.

The improvement also comes with better preservation of general capabilities. We evaluated math on AIME25, code on LiveCodeBench v6, instruction following on IFEval, and science on GPQA Diamond. The average accuracy drop across these benchmarks was 0.9 points for OLIVE, compared with 3.4 points for offline full SFT. Online intervention improved the target task while preserving more performance outside it.

This analysis uses complete, verifier-filtered continuations. The main reasoning comparison above uses short, unfiltered continuations. The plots use the paper’s pass@k and avg@k metric labels.

A text-only teacher keeps the student learning.

OLIVE turns generated teacher text into a continuing source of supervision. On ScienceWorld, we trained Qwen3-1.7B with GPT-5.4-mini as the teacher. Offline SFT reused a fixed set of teacher episodes. OLIVE collected new teacher continuations from the student’s latest interaction histories.

ScienceWorld · Text-only teacher

Fresh interventions sustain learning

Student: Qwen3-1.7B · Teacher: GPT-5.4-mini

Learning across five epochs

Success rate (%)

Both methods train for the number of epochs shown.

OLIVEOffline SFT
Epoch
OLIVE ≈25.8%Offline SFT ≈12.6%
All approximate values
  • Epoch 1: OLIVE ≈8.9%; Offline SFT ≈9.1%.
  • Epoch 2: OLIVE ≈7.8%; Offline SFT ≈13.2%.
  • Epoch 3: OLIVE ≈17.5%; Offline SFT ≈12.9%.
  • Epoch 4: OLIVE ≈20.3%; Offline SFT ≈15.1%.
  • Epoch 5: OLIVE ≈25.8%; Offline SFT ≈12.6%.

More student attempts within one epoch

Success rate (%)

OLIVE: rollouts per prompt in one epoch. Offline SFT: epochs on fixed data.

OLIVEOffline SFT
Rollouts / epochs
OLIVE ≈20.0%Offline SFT ≈15.1%
All approximate values
  • Setting 1: OLIVE ≈8.9%; Offline SFT ≈9.1%.
  • Setting 2: OLIVE ≈14.1%; Offline SFT ≈13.3%.
  • Setting 4: OLIVE ≈20.0%; Offline SFT ≈15.1%.

Curves and error bars are reconstructed from the paper’s plotted coordinates; readouts are approximate. The paper does not specify the error-bar statistic. Both views use the same 0–30% scale.

Offline SFT is ahead during the first two epochs, then its progress levels off. OLIVE keeps improving through the later epochs and finishes well ahead. Refreshing the context changes what the teacher demonstrates as training proceeds, keeping the supervision useful to the evolving student.

More online attempts also help within a single epoch. In the rollout comparison, OLIVE samples multiple prefixes per prompt during one epoch, while the offline baseline revisits its fixed data for more epochs. OLIVE is ahead at the settings with two and four rollouts per prompt. Fresh attempts give the teacher additional student-generated contexts to continue.

Keep learning as tasks change.

We also tested whether online intervention helps a student continue learning across different environments. Starting from Qwen3-4B, we trained for five epochs on each environment in sequence and then evaluated the resulting policy on all five tasks.

Sequential training · Qwen3-4B

Stronger learning across successive tasks

Success-rate difference: OLIVE minus offline distillation.

Training order: BabyAI → TextCraft → SearchQA → ALFWorld → ScienceWorld. Five epochs on each environment.

Gain after sequential training

Percentage points · positive favors OLIVE
  1. BabyAI0.0 pp
  2. TextCraft+8.5 pp
  3. SearchQA−0.6 pp
  4. ALFWorld+4.8 pp
  5. ScienceWorld+13.0 pp
  1. Average+5.1 pp

OLIVE − offline distillation (percentage points)

Success rates evaluated after sequential training on all five environments. The rows follow training order; the separated final row gives the average difference.

OLIVE gains 5.1 percentage points on average. The largest gain is on ScienceWorld, the last environment. BabyAI is tied; SearchQA is 0.6 points lower.

OLIVE finishes ahead on average, with the largest gain on ScienceWorld, the final environment in the sequence. It matches offline distillation on BabyAI and is slightly behind on SearchQA. As earlier training changes the student’s behavior, fresh interventions keep providing demonstrations for the states it now reaches. These results support online supervision as a useful way to sustain learning across tasks.

Larger-scale training is next.

We found online intervention pretty useful. We’re taking OLIVE to larger-scale training soon, and we’re excited to explore its scalability.

Our next experiments will bring OLIVE to larger students and longer training runs. We’re especially interested in how fresh teacher interventions support an increasingly capable student, and how asynchronous generation can keep the training loop efficient as the workload grows.

A teacher’s most useful demonstration can start partway through the learner’s own attempt. OLIVE turns those moments into training data throughout learning. The gains in reasoning and agent performance make us excited to push it further.

About this project

Blog author: Dylan Zhang

Main contributors

Haojin Wang Dylan Zhang (Project Lead)

With collaborators at MIT, MIT-IBM Watson Lab, the University of Pennsylvania (UPenn), and the University of Washington (UW).

Questions or related experiences? Reach me at shizhuo2@illinois.edu.