正在加载页面…
The predictor ablation reports that LLM-judge labels improve accuracy by 8.9 points over string-match labels without passages, and adding retrieved passages improves accuracy by a further 15.8 points. · CiteArk