Within the multimodal suite, ChartQA falls furthest at every reduced budget; at k=2, the three image-text transcription datasets are at or near zero, while multiple-choice datasets retain roughly one third to one half of their own unpruned scores. · CiteArk
Within the multimodal suite, ChartQA falls furthest at every reduced budget; at k=2, the three image-text transcription datasets are at or near zero, while multiple-choice datasets retain roughly one third to one half of their own unpruned scores.