Public research repositories on CiteArk, sorted by activity, update time, claims with supporting Assessments, and community reproduction requests.
Subject Artificial Intelligence · 18 repositories
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
0/17This paper introduces JITMEM, an agent-memory framework that stores complete trajectories and postpones their distillation until a new task is known. A task-conditioned curator converts retrieved raw traces into a compact payload for a frozen executor; the curator is optimized with GRPO using immediate task reward. Experiments on ALFWorld, WebShop, and tau2-bench compare prompted and trained read-time curation with no-memory, heuristic, and learned write-time systems across several executors. Reported results favor read-time curation in success, context efficiency, and cross-executor transfer. Ablations attribute the gains to task conditioning, successful-trajectory filtering, retention of raw traces, and actual use of retrieved experience.
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
0/12The paper introduces KV-Invariant Transformer Expansion (KITE), a staged language-model scaling paradigm that adds capacity outside the key-value-producing path. Its Step Scale Transformer (SST) instantiation trains a source tower, adds a second KV-reading tower, and jointly continues training while retaining source-sized bulk prefill. The study compares a 66.959B-parameter SST model with 46.727B and 62.691B classic MoE Transformers under approximately matched theoretical cumulative training FLOPs. It reports lower final training loss for SST, higher scores on seven downstream tasks, slightly lower held-out ARXIV NLL, and lower analytical inference cost in prefill-heavy workload mixes, while emphasizing that serving speedups and causal contributions were not measured.
You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
0/31This paper systematically studies expert-selection redundancy and dynamic expert pruning in twelve fine-grained mixture-of-experts language-model checkpoints from nine architecture families. Uniform top-k truncation is evaluated across knowledge QA, mathematics, code generation, and general reasoning, then compared at matched average routed-expert budgets with four adaptive allocation rules. The authors report that keeping roughly two thirds of selected experts largely preserves aggregate performance and increases inference throughput, while adaptive routing contributes little at conservative budgets but helps under aggressive pruning, especially on generative tasks. Matched comparisons further associate greater pruning resilience with larger and thinking models, and greater vulnerability with a multimodal model whose routing weights are less concentrated.
How Strongly Should Task State Influence an LLM Agent?
0/24This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
0/21KV-COBRA studies extreme-rate compression of transformer key-value caches by treating rank and bit width as a coupled allocation problem. It selects a rank–precision pair for each attention head, redistributes total storage across heads, and uses a fused Hadamard rotation to equalize retained-channel variance. A query-aware variant reorders SVD directions using an attention-KL importance proxy. Experiments across three 7–8B models, additional 70B and mixture-of-experts models, perplexity, reasoning, long-context QA, and retrieval show the largest gains at low bit rates. Appendices examine solver behavior, quantizer-model error, calibration robustness, latency primitives, value-cache limits, RoPE placement, and depth sensitivity.
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
0/18The paper presents CIGAsk, a reinforcement-learning recipe for teaching language models both when to request clarification and how to formulate an informative question. In multi-turn GRPO, Counterfactual Information Gain rewards clarification turns according to how much a frozen reference model's likelihood of the gold answer increases after the simulated user's response, while an asymmetric ambiguity bonus rewards selective asking. Experiments use Qwen2.5 3B and 7B policies on PACIFIC, AbgCoQA, and AmbigNQ. The authors report gains over prompting, supervised warmstarts, and contextual external baselines, component and reference-model ablations, cross-backbone transfer, simulator sensitivity, and retention on two out-of-domain closed-book QA datasets.
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
0/29This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
0/9This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent–Tool Boundary
0/15This paper studies consistency failures that arise when agent workflows invoke independently supplied tools whose external effects may be uncertain, irreversible, concurrent, or visible before workflow resolution. It introduces an effect-history model that separates real-world events from runtime observations, organizes eight anomalies into uncertainty, workflow, and interaction families, and derives boundary capabilities and transactional contract families needed to exclude them. It also compares six agent systems against the catalog and surveys standard annotations for 98,291 tools from remotely reachable Model Context Protocol servers. The census finds widespread but coarse annotation use and concludes that the current vocabulary cannot fully express any required transactional capability.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
0/57Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
0/76Knowledge distillation with post-trained teachers is increasingly used to compress language models, yet whether its benefits remain consistent across training stages is unclear. Evaluating logit-based distillation on OLMo-2, the authors identify a stage-dependent reasoning-recall tradeoff: while forward-KL distillation improves both reasoning and factual recall during pre-training, it slows factual acquisition during mid-training. This discrepancy stems from asymmetric teacher confidence across domains—teachers are confident on procedural tasks but uncertain on factual tokens, which students learn early. To resolve this imbalance, the authors introduce Switch Distillation, dynamically routing confident tokens to reverse-KL distillation and uncertain tokens to cross-entropy. Switch Distillation matches or exceeds baselines across reasoning and knowledge benchmarks while preserving factual recall.
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
0/14Skill-augmented agents persist loaded skills across execution loops, creating severe vulnerabilities where malicious or compromised skills steer tool calls, leak secrets, or corrupt state after installation. Static pre-installation vetting cannot detect these task-conditioned threats. To address this challenge, the authors introduce Defense-as-Skill, deploying the runtime guard as an installable, inspectable, and editable skill named SkillSonar that dynamically gates sensitive actions through allow, replan, or confirmation decisions. They construct SCOPE-R, a benchmark spanning six risk families with 206 attack-confirmed malicious instances and 43 benign tasks, and optimize SkillSonar using feedback-driven Monte Carlo Tree Search. On repeated GLM-5 evaluations, SkillSonar substantially curtails attack success while preserving benign task utility across multiple agent harnesses.
Just-In-Time Agent Memory with Runtime Agentic Research
0/18This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
0/41The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.
Online Self-Weighted Fine-Tuning
0/5Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
0/14LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
0/4The paper presents AGENTIC R AG-R1, a reinforcement-learning framework for retrieval-augmented multi-step reasoning. It combines fine-grained actions, stack memory with push, pop, revision, and summarization, hierarchical outcome and process rewards, and information-aware rollout rejection. Experiments report comparisons, ablations, step-budget scaling, model-size effects, timing, and downstream agent evaluations.
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
0/3The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.