Loading page…
The predictor ablation reports that LLM-judge labels improve accuracy by 8.9 points over string-match labels without passages, and adding retrieved passages improves accuracy by a further 15.8 points. · CiteArk