arXiv preprint 2026
LongPIBench: A Long-Context Benchmark for Prompt Injection
citeark/longpibench-a-long-context-benchmark-for-prompt-injection--6rf5kr
LongPIBench introduces a benchmark for prompt-injection attacks and defenses in four document-centric long-context workflows: paper peer review, resume screening, email summarization, and code review. It combines synthetic and real-world datasets, evaluates heuristic and optimization-based attacks across eight language models, and measures both prevention-based attack success and detection-based false-positive and false-negative rates. The paper reports that attacks remain effective and that defenses that perform well on short-context benchmarks often degrade on long documents, with additional ablations over document format, injection position, attack goals, context length, and segmentation.
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
citeark/finding-where-the-buck-stops-an-automated-failure-attribution-based-reflection-f--j878vv
The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
citeark/just-in-time-memory-learning-to-curate-task-adaptive-memory-for-llm-agents--15g4ws
This paper introduces JITMEM, an agent-memory framework that stores complete trajectories and postpones their distillation until a new task is known. A task-conditioned curator converts retrieved raw traces into a compact payload for a frozen executor; the curator is optimized with GRPO using immediate task reward. Experiments on ALFWorld, WebShop, and tau2-bench compare prompted and trained read-time curation with no-memory, heuristic, and learned write-time systems across several executors. Reported results favor read-time curation in success, context efficiency, and cross-executor transfer. Ablations attribute the gains to task conditioning, successful-trajectory filtering, retention of raw traces, and actual use of retrieved experience.
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
citeark/kite-kv-invariant-transformer-expansion-for-efficient-agentic-llm-scaling--8nfw3u
The paper introduces KV-Invariant Transformer Expansion (KITE), a staged language-model scaling paradigm that adds capacity outside the key-value-producing path. Its Step Scale Transformer (SST) instantiation trains a source tower, adds a second KV-reading tower, and jointly continues training while retaining source-sized bulk prefill. The study compares a 66.959B-parameter SST model with 46.727B and 62.691B classic MoE Transformers under approximately matched theoretical cumulative training FLOPs. It reports lower final training loss for SST, higher scores on seven downstream tasks, slightly lower held-out ARXIV NLL, and lower analytical inference cost in prefill-heavy workload mixes, while emphasizing that serving speedups and causal contributions were not measured.
Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models
citeark/informed-masking-structure-aware-perturbation-for-reinforcement-learning-in-diff--94axkp
This paper studies how masked reconstruction subproblems are selected when reinforcement-learning objectives for diffusion language models can sample only a few perturbations per rollout. It identifies an upstream/downstream structure from confidence changes along denoising trajectories and defines a token priority score. Informed Masking uses that score to bias perturbations toward downstream tokens while leaving the host method’s reward, mask budget, and estimator intact. Experiments with LLaDA-8B-Instruct across GSM8K, MATH-500, Countdown, and Sudoku integrate the method into ESPO, GDPO, and SPG. Reported results show generally higher accuracy, better-conditioned reconstruction estimates, reduced gradient dispersion, faster convergence, and improved stability, alongside limitations in attribution and decoding-order dependence.
You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
citeark/you-only-need-2-3-of-the-chosen-experts-an-empirical-study-of-dynamic-expert-pru--a34ppx
This paper systematically studies expert-selection redundancy and dynamic expert pruning in twelve fine-grained mixture-of-experts language-model checkpoints from nine architecture families. Uniform top-k truncation is evaluated across knowledge QA, mathematics, code generation, and general reasoning, then compared at matched average routed-expert budgets with four adaptive allocation rules. The authors report that keeping roughly two thirds of selected experts largely preserves aggregate performance and increases inference throughput, while adaptive routing contributes little at conservative budgets but helps under aggressive pruning, especially on generative tasks. Matched comparisons further associate greater pruning resilience with larger and thinking models, and greater vulnerability with a multimodal model whose routing weights are less concentrated.
How Strongly Should Task State Influence an LLM Agent?
citeark/how-strongly-should-task-state-influence-an-llm-agent--8zd84a
This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
citeark/kv-cobra-kv-cache-compression-via-co-optimized-bit-rank-allocation--0beowq
KV-COBRA studies extreme-rate compression of transformer key-value caches by treating rank and bit width as a coupled allocation problem. It selects a rank–precision pair for each attention head, redistributes total storage across heads, and uses a fused Hadamard rotation to equalize retained-channel variance. A query-aware variant reorders SVD directions using an attention-KL importance proxy. Experiments across three 7–8B models, additional 70B and mixture-of-experts models, perplexity, reasoning, long-context QA, and retrieval show the largest gains at low bit rates. Appendices examine solver behavior, quantizer-model error, calibration robustness, latency primitives, value-cache limits, RoPE placement, and depth sensitivity.
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
citeark/when-and-how-should-an-agent-clarify-cigask-teaching-llms-to-clarify-via-counter--iisl0l
The paper presents CIGAsk, a reinforcement-learning recipe for teaching language models both when to request clarification and how to formulate an informative question. In multi-turn GRPO, Counterfactual Information Gain rewards clarification turns according to how much a frozen reference model's likelihood of the gold answer increases after the simulated user's response, while an asymmetric ambiguity bonus rewards selective asking. Experiments use Qwen2.5 3B and 7B policies on PACIFIC, AbgCoQA, and AmbigNQ. The authors report gains over prompting, supervised warmstarts, and contextual external baselines, component and reference-model ablations, cross-backbone transfer, simulator sensitivity, and retention on two out-of-domain closed-book QA datasets.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
citeark/abstention-and-noise-filtering-two-missing-primitives-of-softmax-attention--jien37
This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
citeark/aspire-asynchronous-batched-self-speculative-decoding-for-long-context-llm-infer--f43ppe
ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
citeark/diagnosing-the-fact-grounding-gap-in-multi-hop-question-answering--arkkn9
This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
citeark/paperdoctor-evidence-grounded-and-actionable-feedback-for-scientific-papers-in-p--xmpjzq
PaperDoctor is an agent framework for diagnosing draft scientific papers rather than issuing acceptance verdicts. It parses manuscripts and code, performs low-cost writing, visual, citation, and claim screening, routes extracted claims to code, theory, literature, and experiment-design verifiers, and selectively reruns experiments. Each finding links an observation to source evidence and a proposed revision. The paper evaluates author responses on 30 in-progress papers and compares PaperDoctor with referees and an agentic reviewer on 40 papers across four domains. It also analyzes experiment reproduction, showing both frequent pre-execution blockers and substantial post-execution mismatches.
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
citeark/beyond-token-local-imitation-reward-compatible-temporal-credit-assignment-for-on--7jye4o
This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
citeark/nameless-tokenization-a-lossless-tokenizer-level-defense-against-control-token-f--p9pmav
Open-weight chat models expose the surface strings that their tokenizers map to reserved turn, role, reasoning, and tool identifiers, allowing attacker-controlled text to forge structural tokens. The paper audits 256 deployed chat tokenizers, evaluates a tokenizer-only defense called nameless tokenization on five model families, and compares it with string sanitizers and training-level instruction–data separation. Nameless tokenization removes text-to-control-identifier mappings while preserving template-written identifiers and message characters. Reported experiments cover attack-free utility, delimiter-bearing content, several forged-turn attacks, two injected objectives, two system-prompt conditions, and model-specific outcomes. Results show exact attack-free token-stream preservation, much better delimiter fidelity, and strongly context-dependent security gains.
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent–Tool Boundary
citeark/when-tool-calls-succeed-but-workflows-fail-anomalies-at-the-agent-tool-boundary--so6e7v
This paper studies consistency failures that arise when agent workflows invoke independently supplied tools whose external effects may be uncertain, irreversible, concurrent, or visible before workflow resolution. It introduces an effect-history model that separates real-world events from runtime observations, organizes eight anomalies into uncertainty, workflow, and interaction families, and derives boundary capabilities and transactional contract families needed to exclude them. It also compares six agent systems against the catalog and surveys standard annotations for 98,291 tools from remotely reachable Model Context Protocol servers. The census finds widespread but coarse annotation use and concludes that the current vocabulary cannot fully express any required transactional capability.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
citeark/procedural-graphs-self-evolving-execution-structures-for-llm-agents--cbtk6h
Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
citeark/knowledge-distillation-during-mid-training-favors-reasoning-over-factual-recall--dt61dt
Knowledge distillation with post-trained teachers is increasingly used to compress language models, yet whether its benefits remain consistent across training stages is unclear. Evaluating logit-based distillation on OLMo-2, the authors identify a stage-dependent reasoning-recall tradeoff: while forward-KL distillation improves both reasoning and factual recall during pre-training, it slows factual acquisition during mid-training. This discrepancy stems from asymmetric teacher confidence across domains—teachers are confident on procedural tasks but uncertain on factual tokens, which students learn early. To resolve this imbalance, the authors introduce Switch Distillation, dynamically routing confident tokens to reverse-KL distillation and uncertain tokens to cross-entropy. Switch Distillation matches or exceeds baselines across reasoning and knowledge benchmarks while preserving factual recall.
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
citeark/defense-as-skill-evolving-runtime-guard-skill-for-skill-augmented-agents--rl2sck
Skill-augmented agents persist loaded skills across execution loops, creating severe vulnerabilities where malicious or compromised skills steer tool calls, leak secrets, or corrupt state after installation. Static pre-installation vetting cannot detect these task-conditioned threats. To address this challenge, the authors introduce Defense-as-Skill, deploying the runtime guard as an installable, inspectable, and editable skill named SkillSonar that dynamically gates sensitive actions through allow, replan, or confirmation decisions. They construct SCOPE-R, a benchmark spanning six risk families with 206 attack-confirmed malicious instances and 43 benign tasks, and optimize SkillSonar using feedback-driven Monte Carlo Tree Search. On repeated GLM-5 evaluations, SkillSonar substantially curtails attack success while preserving benign task utility across multiple agent harnesses.
Generalization Dynamics of LM Pre-training
citeark/generalization-dynamics-of-lm-pre-training--awr6et
This study tracks language-model generalization during pretraining using six behavioral evaluations and additional fine-tuning tests. Across OLMo3 and Apertus checkpoints, models repeatedly switch between following shallow patterns and producing answers consistent with task structure, even late in training. The reported fluctuations persist with probability-based metrics and alternative training-compute axes, while common benchmarks remain comparatively stable. Single-step updates and checkpoint averaging do not eliminate the phenomenon. Larger models generalize more often, but still fluctuate. Intermediate checkpoint selection improves post-training reasoning transfer and resistance to prefilling attacks in the reported comparisons. A preliminary data-selection experiment also steers generalization dynamics. Complexity proxies give inconsistent predictions across layers and datasets.
Just-In-Time Agent Memory with Runtime Agentic Research
citeark/just-in-time-agent-memory-with-runtime-agentic-research--0lc4uw
This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.
Online Self-Weighted Fine-Tuning
citeark/online-self-weighted-fine-tuning--iqyzxx
Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
citeark/geometry-of-divergence-tracking-hidden-state-trajectories-for-adaptive-multi-tur--z06w8y
LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
citeark/agenticrag-r1-agentic-reinforcement-learning-with-stack-memory-for-multi-step-re--drrsza
The paper presents AGENTIC R AG-R1, a reinforcement-learning framework for retrieval-augmented multi-step reasoning. It combines fine-grained actions, stack memory with push, pop, revision, and summarization, hierarchical outcome and process rewards, and information-aware rollout rejection. Experiments report comparisons, ablations, step-budget scaling, model-size effects, timing, and downstream agent evaluations.
LoGo: Token-Level Dynamic Local-Global Attention
citeark/logo-token-level-dynamic-local-global-attention--k7rv0t
The paper introduces LoGo, a decoder-only Transformer attention mechanism that gives every token a local causal-attention branch and selectively activates a full-context branch through a learned token-level gate. An adaptive threshold controls the global activation ratio without an auxiliary balancing loss, progressive masking stabilizes training, and query-sparse Triton kernels avoid computing global attention for unselected queries. Experiments compare LoGo with full-attention and static local-global hybrids across model scales, long-context extension stages, recall benchmarks, kernel runtimes, ablations, and routing analyses.