Performance on ten common datasets remains comparatively stable across pretraining in both model families, unlike the toy generalization probes. Standard sentiment/topic accuracy is strong; other benchmark performance often improves gradually. · CiteArk
Performance on ten common datasets remains comparatively stable across pretraining in both model families, unlike the toy generalization probes. Standard sentiment/topic accuracy is strong; other benchmark performance often improves gradually.