1 experiments
Execution agent’s result notes
Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance).
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).
Metric: roc_auc · score
Missing values are not zero.
Data and sources · 16
| Source | Conditions | Metric | Value | Assessment and limits |
|---|---|---|---|---|
| Papersource_paper:Figure 8, Page 27, Cell 1b baseclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_base/paper | — | roc_auc | 0.815 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 1b baseclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_base/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | 0.8079 score | supported |
| Papersource_paper:Figure 8, Page 27, Cell 1b sftclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_sft/paper | — | roc_auc | 0.77 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 1b sftclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_sft/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | 0.7967 score | supported |
| Papersource_paper:Figure 8, Page 27, Cell 1b dpoclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_dpo/paper | — | roc_auc | 0.744 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 1b dpoclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_dpo/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | 0.7931 score | supported |
| Papersource_paper:Figure 8, Page 27, Cell 1b instructclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_instruct/paper | — | roc_auc | 0.748 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 1b instructclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_1b_instruct/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | 0.7973 score | supported |
| Papersource_paper:Figure 8, Page 27, Cell 7b baseclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_base/paper | — | roc_auc | 0.816 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 7b baseclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_base/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | — score | inconclusive |
| Papersource_paper:Figure 8, Page 27, Cell 7b sftclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_sft/paper | — | roc_auc | 0.777 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 7b sftclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_sft/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | — score | inconclusive |
| Papersource_paper:Figure 8, Page 27, Cell 7b dpoclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_dpo/paper | — | roc_auc | 0.76 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 7b dpoclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_dpo/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | — score | inconclusive |
| Papersource_paper:Figure 8, Page 27, Cell 7b instructclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_instruct/paper | — | roc_auc | 0.761 score | reported |
| Reproductionsource_paper:Figure 8, Page 27, Cell 7b instructclaim_fig8_asymmetric_supervision_stages/fig8_roc_auc_7b_instruct/urn:citeark:assessment:35681bd52a61526edce23488b64679c9cb67d3a816856d673d613ef2f550b487 | — | roc_auc | — score | inconclusive |
Comparison conditions
Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models).
Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.
Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json
Evaluated only OLMo-2 1B checkpoints; all 8 measurements for OLMo-2 7B and 13B models across Base, SFT, DPO, and Instruct stages were omitted due to checkpoint download sizes (>220 GB total) and budget limits.
The 240 Dolmino evaluation documents were reconstructed by sampling and token adjustment from Dolmino Mix 1124 rather than using the exact fixed author test split/seed.
The claim posits that teacher predictive entropy separates procedural from knowledge-intensive domains with high ROC AUC across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B). However, the executed reproduction only evaluated the 1B model checkpoints due to download size and runtime budget constraints. The 8 measurements for 7B and 13B models across all stages were not evaluated, leaving the claim coverage partial. Furthermore, the 240 documents used were constructed via synthetic token adjustment/sampling rather than the exact fixed evaluation subset from the paper, introducing minor protocol divergence. While the 1B measurements were successfully executed and reproduced high ROC AUC values (~0.79-0.81), the entire claim remains inconclusive due to missing multi-model coverage. Some original claim measurements remain without evidence; available measurements are assessed individually.