Performance on ten common datasets remains comparatively stable across pretraining in both model families, unlike the toy generalization probes. Standard sentiment/topic accuracy is strong; other benchmark performance often improves gradually. · CiteArk
Not assessedPlan blockedFindingordinary-benchmarks-stable
Performance on ten common datasets remains comparatively stable across pretraining in both model families, unlike the toy generalization probes. Standard sentiment/topic accuracy is strong; other benchmark performance often improves gradually.