学科 多智能体系统 · 3 个仓库
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
0/21PaperDoctor is an agent framework for diagnosing draft scientific papers rather than issuing acceptance verdicts. It parses manuscripts and code, performs low-cost writing, visual, citation, and claim screening, routes extracted claims to code, theory, literature, and experiment-design verifiers, and selectively reruns experiments. Each finding links an observation to source evidence and a proposed revision. The paper evaluates author responses on 30 in-progress papers and compares PaperDoctor with referees and an agentic reviewer on 40 papers across four domains. It also analyzes experiment reproduction, showing both frequent pre-execution blockers and substantial post-execution mismatches.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
0/57Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
0/4The paper presents AGENTIC R AG-R1, a reinforcement-learning framework for retrieval-augmented multi-step reasoning. It combines fine-grained actions, stack memory with push, pop, revision, and summarization, hierarchical outcome and process rewards, and information-aware rollout rejection. Experiments report comparisons, ablations, step-budget scaling, model-size effects, timing, and downstream agent evaluations.