Loading page…
Resource use and parsing robustness across PG construction modes on MultiChallenge (56-sample test split) with Gemini 3.5 Flash. Mode 5 (Scratch + Online Evolution) achieves a 45.7% reduction in parsing failures per sample relative to unguided baseline (0.57 vs 1.05), while using 5.05 steps and 92.81s per sample (compared to 8.02 steps and 166.57s for Mode 4). Mode 3 achieves highest overall success (92.86%) at 6.73 steps and 128.50s. · CiteArk