Across late-training checkpoints, PACIFIC items that the policy chose to clarify had higher F1 than items answered directly, with bootstrap 95% confidence intervals for the difference above zero throughout. · CiteArk
Not assessedPlan blockedFindinglate-checkpoint-productivity-001
Across late-training checkpoints, PACIFIC items that the policy chose to clarify had higher F1 than items answered directly, with bootstrap 95% confidence intervals for the difference above zero throughout.
Source: paper:PDF p. 7, Section 5.1
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.