Reproduction materials
Prepared 0/0 materials
资料与依赖由同一实验工作区准备,直接复用已有缓存,不另启准备 Agent。
This experiment has no separate content download, but dependencies, configuration, and loaders are still checked.
Prefetching only saves time. Unprepared material is handed to the complete reproduction Agent to download, adapt, and verify autonomously.
Credit estimate for this reproduction
Expected usage
509
Maximum reservation
1323
Estimated compute time
About 195 min
Actual usage was 156 credits; 0 credits were charged. Compute ran for about 9 minutes and the model used 6209368 tokens.
Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
Supervised fine-tuning of large language models on reasoning tasks traditionally applies equal gradient updates to every example, wasting capacity and injecting gradient noise on already-solved problems. OSW-FT bridges supervised fine-tuning and reinforcement learning by using lightweight online rollouts to estimate an empirical failure rate that dynamically scales the loss on expert trajectories. This allows post-training practitioners to focus gradient updates on unsolved capability frontiers without incurring the high sample complexity and training instability of full policy-gradient reinforcement learning. The primary limitation is its dependence on binary-verifiable reward verifiers and available high-quality expert demonstrations, restricting direct applicability in open-ended or non-verifiable domains.
The paper’s claims are available in Research claims.