Public research repositories on CiteArk, sorted by activity, update time, claims with supporting Assessments, and community reproduction requests.
Subject Computation and Language · 16 repositories
Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models
0/18This paper studies how masked reconstruction subproblems are selected when reinforcement-learning objectives for diffusion language models can sample only a few perturbations per rollout. It identifies an upstream/downstream structure from confidence changes along denoising trajectories and defines a token priority score. Informed Masking uses that score to bias perturbations toward downstream tokens while leaving the host method’s reward, mask budget, and estimator intact. Experiments with LLaDA-8B-Instruct across GSM8K, MATH-500, Countdown, and Sudoku integrate the method into ESPO, GDPO, and SPG. Reported results show generally higher accuracy, better-conditioned reconstruction estimates, reduced gradient dispersion, faster convergence, and improved stability, alongside limitations in attribution and decoding-order dependence.
How Strongly Should Task State Influence an LLM Agent?
0/24This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
0/16This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
0/23ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
0/29This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
0/21PaperDoctor is an agent framework for diagnosing draft scientific papers rather than issuing acceptance verdicts. It parses manuscripts and code, performs low-cost writing, visual, citation, and claim screening, routes extracted claims to code, theory, literature, and experiment-design verifiers, and selectively reruns experiments. Each finding links an observation to source evidence and a proposed revision. The paper evaluates author responses on 30 in-progress papers and compares PaperDoctor with referees and an agentic reviewer on 40 papers across four domains. It also analyzes experiment reproduction, showing both frequent pre-execution blockers and substantial post-execution mismatches.
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
0/13Open-weight chat models expose the surface strings that their tokenizers map to reserved turn, role, reasoning, and tool identifiers, allowing attacker-controlled text to forge structural tokens. The paper audits 256 deployed chat tokenizers, evaluates a tokenizer-only defense called nameless tokenization on five model families, and compares it with string sanitizers and training-level instruction–data separation. Nameless tokenization removes text-to-control-identifier mappings while preserving template-written identifiers and message characters. Reported experiments cover attack-free utility, delimiter-bearing content, several forged-turn attacks, two injected objectives, two system-prompt conditions, and model-specific outcomes. Results show exact attack-free token-stream preservation, much better delimiter fidelity, and strongly context-dependent security gains.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
0/57Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
0/76Knowledge distillation with post-trained teachers is increasingly used to compress language models, yet whether its benefits remain consistent across training stages is unclear. Evaluating logit-based distillation on OLMo-2, the authors identify a stage-dependent reasoning-recall tradeoff: while forward-KL distillation improves both reasoning and factual recall during pre-training, it slows factual acquisition during mid-training. This discrepancy stems from asymmetric teacher confidence across domains—teachers are confident on procedural tasks but uncertain on factual tokens, which students learn early. To resolve this imbalance, the authors introduce Switch Distillation, dynamically routing confident tokens to reverse-KL distillation and uncertain tokens to cross-entropy. Switch Distillation matches or exceeds baselines across reasoning and knowledge benchmarks while preserving factual recall.
Generalization Dynamics of LM Pre-training
0/21This study tracks language-model generalization during pretraining using six behavioral evaluations and additional fine-tuning tests. Across OLMo3 and Apertus checkpoints, models repeatedly switch between following shallow patterns and producing answers consistent with task structure, even late in training. The reported fluctuations persist with probability-based metrics and alternative training-compute axes, while common benchmarks remain comparatively stable. Single-step updates and checkpoint averaging do not eliminate the phenomenon. Larger models generalize more often, but still fluctuate. Intermediate checkpoint selection improves post-training reasoning transfer and resistance to prefilling attacks in the reported comparisons. A preliminary data-selection experiment also steers generalization dynamics. Complexity proxies give inconsistent predictions across layers and datasets.
Just-In-Time Agent Memory with Runtime Agentic Research
0/18This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
0/41The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.
LongPIBench: A Long-Context Benchmark for Prompt Injection
0/4LongPIBench introduces a benchmark for prompt-injection attacks and defenses in four document-centric long-context workflows: paper peer review, resume screening, email summarization, and code review. It combines synthetic and real-world datasets, evaluates heuristic and optimization-based attacks across eight language models, and measures both prevention-based attack success and detection-based false-positive and false-negative rates. The paper reports that attacks remain effective and that defenses that perform well on short-context benchmarks often degrade on long documents, with additional ablations over document format, injection position, attack goals, context length, and segmentation.
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
0/14LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
LoGo: Token-Level Dynamic Local-Global Attention
0/6The paper introduces LoGo, a decoder-only Transformer attention mechanism that gives every token a local causal-attention branch and selectively activates a full-context branch through a learned token-level gate. An adaptive threshold controls the global activation ratio without an auxiliary balancing loss, progressive masking stabilizes training, and query-sparse Triton kernels avoid computing global attention for unselected queries. Experiments compare LoGo with full-attention and static local-global hybrids across model scales, long-context extension stages, recall benchmarks, kernel runtimes, ablations, and routing analyses.
Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
0/3The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.