Loading page…
Across 24 model-benchmark evaluation settings (4 LLM families x 6 benchmarks), Procedural Graph ranks first or joint first in 21 of 24 settings. Compared against the strongest baseline in each setting, PG records 19 wins, 2 ties, and 3 losses (one-sided exact binomial sign test excluding ties, p = 4.3e-4). · CiteArk