复现资料准备
已准备 0/0 项资料
资料与依赖由同一实验工作区准备,直接复用已有缓存,不另启准备 Agent。
本实验没有需要单独下载的正文资产,仍会核验依赖、配置和加载器。
预下载只是节省时间的优化;未准备的内容已交给完整复现 Agent 自主下载、调整和验证。
本次复现积分预估
预计消耗
509
最大预留
1323
预计算力时长
约 195 分钟
实际消耗 156 积分,扣除 0 积分;算力运行约 9 分钟,模型 Token 6209368。
Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
Supervised fine-tuning of large language models on reasoning tasks traditionally applies equal gradient updates to every example, wasting capacity and injecting gradient noise on already-solved problems. OSW-FT bridges supervised fine-tuning and reinforcement learning by using lightweight online rollouts to estimate an empirical failure rate that dynamically scales the loss on expert trajectories. This allows post-training practitioners to focus gradient updates on unsolved capability frontiers without incurring the high sample complexity and training instability of full policy-gradient reinforcement learning. The primary limitation is its dependence on binary-verifiable reward verifiers and available high-quality expert demonstrations, restricting direct applicability in open-ended or non-verifiable domains.
论文中的结论可在「研究结论」中查看。