Step-level generative guidance introduces substantial token overhead relative to unguided baselines, even when it shortens solver trajectories. On GDPval and ALFWorld, localized guidance reduces average solver steps from 28.20 to 18.57 and from 21.84 to 18.80, respectively, but total token consumption remains 33.4% and 55.4% higher due to the per-step guidance model call. · CiteArk
Step-level generative guidance introduces substantial token overhead relative to unguided baselines, even when it shortens solver trajectories. On GDPval and ALFWorld, localized guidance reduces average solver steps from 28.20 to 18.57 and from 21.84 to 18.80, respectively, but total token consumption remains 33.4% and 55.4% higher due to the per-step guidance model call.