Loading page…
The study uses only GPT-4.1-mini as the reasoning model, so failure rates and intervention effects may differ for other reasoning models despite cross-dataset and cross-retriever results. · CiteArk