Increasing the classic baseline from 47B to 63B does not improve every task: the 63B classic score is lower on GSM8K (55.04 versus 56.56) and HumanEval few-shot (32.93 versus 35.98), although it is higher on the other reported evaluation columns. · CiteArk
Not assessedPlan blockedFindingclm-classic-scaling-nonmonotonic
Increasing the classic baseline from 47B to 63B does not improve every task: the 63B classic score is lower on GSM8K (55.04 versus 56.56) and HumanEval few-shot (32.93 versus 35.98), although it is higher on the other reported evaluation columns.
Source: fixed-paper-pdf:PDF p. 7, Table 3
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.