正在加载页面…
The paper reports strong agreement between GPT-4.1-mini judging and human annotation, while noting a conservative error skew for fact-presence decisions. · CiteArk