Reproduction materials
Prepared 0/1 materials
科研资料尚未闭合:configuration。这些问题只影响预下载缓存,不会阻止最终复现 Agent。
The preparation method is selected after link checks and inventory review.
Credit estimate for this reproduction
Expected usage
127
Maximum reservation
508
Estimated compute time
About 5 min
Actual usage was 475 credits; 0 credits were charged. Compute ran for about 22 minutes and the model used 18963722 tokens.
Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
Procedural Graphs shift the burden of procedural coherence from unconstrained model generation into an explicit, editable, and inspectable external graph without model parameter fine-tuning. This framework enables agents to navigate complex multi-step tasks, avoid known pitfalls, and improve autonomously through closed-loop trajectory feedback. However, per-step generative guidance introduces substantial token overhead, and evolutionary refinement decisions on small validation sets can be sensitive to episode variance.
The paper’s claims are available in Research claims.