33 个仓库
Simplifying Graph Convolutional Networks
1/27This paper derives Simple Graph Convolution (SGC) by removing intermediate nonlinearities from a graph convolutional network and collapsing its weight matrices into one linear classifier. The resulting method applies a fixed, parameter-free graph propagation filter to node features before multinomial logistic regression. A spectral analysis relates propagation with self-loops to low-pass filtering and proves that self-loops shrink the normalized-Laplacian spectrum. Experiments compare SGC with graph neural-network baselines on citation and social networks and adapt it to text classification, geolocation, relation extraction, zero-shot image classification, graph classification, and molecular prediction, emphasizing accuracy, training time, stability, and known failure cases.
Random Features for Large-Scale Kernel Machines
0/13Rahimi and Recht introduce randomized, explicit low-dimensional feature maps whose Euclidean inner products approximate shift-invariant kernels, allowing nonlinear kernel methods to be replaced by fast linear learning. They develop random Fourier features from a kernel's spectral distribution and random binning features from randomly shifted grids, and give uniform approximation bounds for both constructions. Using ridge regression on five large-scale regression and classification datasets, they compare the two feature families with Core Vector Machines and published exact-kernel baselines. Reported results show competitive test error with substantial, though dataset-dependent, training-time advantages, while also revealing differences between interpolation-oriented Fourier features and locality-preserving binning features.
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
0/17This paper introduces JITMEM, an agent-memory framework that stores complete trajectories and postpones their distillation until a new task is known. A task-conditioned curator converts retrieved raw traces into a compact payload for a frozen executor; the curator is optimized with GRPO using immediate task reward. Experiments on ALFWorld, WebShop, and tau2-bench compare prompted and trained read-time curation with no-memory, heuristic, and learned write-time systems across several executors. Reported results favor read-time curation in success, context efficiency, and cross-executor transfer. Ablations attribute the gains to task conditioning, successful-trajectory filtering, retention of raw traces, and actual use of retrieved experience.
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
0/12The paper introduces KV-Invariant Transformer Expansion (KITE), a staged language-model scaling paradigm that adds capacity outside the key-value-producing path. Its Step Scale Transformer (SST) instantiation trains a source tower, adds a second KV-reading tower, and jointly continues training while retaining source-sized bulk prefill. The study compares a 66.959B-parameter SST model with 46.727B and 62.691B classic MoE Transformers under approximately matched theoretical cumulative training FLOPs. It reports lower final training loss for SST, higher scores on seven downstream tasks, slightly lower held-out ARXIV NLL, and lower analytical inference cost in prefill-heavy workload mixes, while emphasizing that serving speedups and causal contributions were not measured.
Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models
0/18This paper studies how masked reconstruction subproblems are selected when reinforcement-learning objectives for diffusion language models can sample only a few perturbations per rollout. It identifies an upstream/downstream structure from confidence changes along denoising trajectories and defines a token priority score. Informed Masking uses that score to bias perturbations toward downstream tokens while leaving the host method’s reward, mask budget, and estimator intact. Experiments with LLaDA-8B-Instruct across GSM8K, MATH-500, Countdown, and Sudoku integrate the method into ESPO, GDPO, and SPG. Reported results show generally higher accuracy, better-conditioned reconstruction estimates, reduced gradient dispersion, faster convergence, and improved stability, alongside limitations in attribution and decoding-order dependence.
You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
0/31This paper systematically studies expert-selection redundancy and dynamic expert pruning in twelve fine-grained mixture-of-experts language-model checkpoints from nine architecture families. Uniform top-k truncation is evaluated across knowledge QA, mathematics, code generation, and general reasoning, then compared at matched average routed-expert budgets with four adaptive allocation rules. The authors report that keeping roughly two thirds of selected experts largely preserves aggregate performance and increases inference throughput, while adaptive routing contributes little at conservative budgets but helps under aggressive pruning, especially on generative tasks. Matched comparisons further associate greater pruning resilience with larger and thinking models, and greater vulnerability with a multimodal model whose routing weights are less concentrated.
How Strongly Should Task State Influence an LLM Agent?
0/24This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
0/21KV-COBRA studies extreme-rate compression of transformer key-value caches by treating rank and bit width as a coupled allocation problem. It selects a rank–precision pair for each attention head, redistributes total storage across heads, and uses a fused Hadamard rotation to equalize retained-channel variance. A query-aware variant reorders SVD directions using an attention-KL importance proxy. Experiments across three 7–8B models, additional 70B and mixture-of-experts models, perplexity, reasoning, long-context QA, and retrieval show the largest gains at low bit rates. Appendices examine solver behavior, quantizer-model error, calibration robustness, latency primitives, value-cache limits, RoPE placement, and depth sensitivity.
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
0/18The paper presents CIGAsk, a reinforcement-learning recipe for teaching language models both when to request clarification and how to formulate an informative question. In multi-turn GRPO, Counterfactual Information Gain rewards clarification turns according to how much a frozen reference model's likelihood of the gold answer increases after the simulated user's response, while an asymmetric ambiguity bonus rewards selective asking. Experiments use Qwen2.5 3B and 7B policies on PACIFIC, AbgCoQA, and AmbigNQ. The authors report gains over prompting, supervised warmstarts, and contextual external baselines, component and reference-model ablations, cross-backbone transfer, simulator sensitivity, and retention on two out-of-domain closed-book QA datasets.
Auto-Encoding Variational Bayes
2/13This paper develops stochastic-gradient variational inference for directed probabilistic models with continuous latent variables and intractable posteriors. Its reparameterization expresses posterior samples as differentiable transformations of parameter-free noise, yielding the Stochastic Gradient Variational Bayes estimator. For independent observations with local latent variables, the Auto-Encoding Variational Bayes algorithm jointly trains a probabilistic encoder and generative decoder, producing the variational auto-encoder when both are neural networks. Experiments on MNIST and Frey Face compare AEVB with wake-sleep using variational lower bounds, and on low-dimensional MNIST also compare estimated marginal likelihood with Monte Carlo EM. Appendix visualizations show learned manifolds and generated samples.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
0/16This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
0/23ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
0/29This paper studies whether passages retrieved during multi-hop question answering contain the specific relational facts required at each reasoning step. Using a Self-Ask pipeline across MuSiQue, HotpotQA, and 2WikiMultihopQA, it distinguishes retrieval failures from extraction failures, where a gold passage is retrieved but the required fact remains unavailable. It validates an LLM fact-presence judge, trains a DeBERTa predictor of deficient hops, and evaluates targeted augmentation and reranking. The reported results show substantial extraction failures, weak diagnostic value from retrieval recall alone, selective benefits from re-retrieval on retrieval failures, and pronounced variation across datasets and question types.
Deep Residual Learning for Image Recognition
0/22This paper introduces residual learning for training substantially deeper neural networks. Instead of directly fitting a desired mapping, stacked layers learn a residual that is added to an identity shortcut. The authors compare plain and residual networks on ImageNet and CIFAR-10, examine shortcut variants and residual-response magnitudes, and scale residual networks to 152 layers on ImageNet and 1202 layers on CIFAR-10. Reported results show reduced optimization degradation and improved classification accuracy with depth. The learned representations also improve Faster R-CNN detection on PASCAL VOC and COCO and support competitive ImageNet detection and localization systems.
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
0/21PaperDoctor is an agent framework for diagnosing draft scientific papers rather than issuing acceptance verdicts. It parses manuscripts and code, performs low-cost writing, visual, citation, and claim screening, routes extracted claims to code, theory, literature, and experiment-design verifiers, and selectively reruns experiments. Each finding links an observation to source evidence and a proposed revision. The paper evaluates author responses on 30 in-progress papers and compares PaperDoctor with referees and an agentic reviewer on 40 papers across four domains. It also analyzes experiment reproduction, showing both frequent pre-execution blockers and substantial post-execution mismatches.
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
0/9This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
0/13Open-weight chat models expose the surface strings that their tokenizers map to reserved turn, role, reasoning, and tool identifiers, allowing attacker-controlled text to forge structural tokens. The paper audits 256 deployed chat tokenizers, evaluates a tokenizer-only defense called nameless tokenization on five model families, and compares it with string sanitizers and training-level instruction–data separation. Nameless tokenization removes text-to-control-identifier mappings while preserving template-written identifiers and message characters. Reported experiments cover attack-free utility, delimiter-bearing content, several forged-turn attacks, two injected objectives, two system-prompt conditions, and model-specific outcomes. Results show exact attack-free token-stream preservation, much better delimiter fidelity, and strongly context-dependent security gains.
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent–Tool Boundary
0/15This paper studies consistency failures that arise when agent workflows invoke independently supplied tools whose external effects may be uncertain, irreversible, concurrent, or visible before workflow resolution. It introduces an effect-history model that separates real-world events from runtime observations, organizes eight anomalies into uncertainty, workflow, and interaction families, and derives boundary capabilities and transactional contract families needed to exclude them. It also compares six agent systems against the catalog and surveys standard annotations for 98,291 tools from remotely reachable Model Context Protocol servers. The census finds widespread but coarse annotation use and concludes that the current vocabulary cannot fully express any required transactional capability.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
0/27FlashAttention-2 is an exact GPU attention algorithm and implementation that improves on FlashAttention by reducing non-matrix-multiplication work, parallelizing computation across sequence blocks, and repartitioning work among warps to reduce shared-memory communication. The paper benchmarks forward, backward, and combined attention on A100 GPUs across causal and non-causal settings, two head dimensions, and sequence lengths from 512 to 16k, and also reports unoptimized H100 results. It evaluates end-to-end GPT-style training for 1.3B- and 2.7B-parameter models at 2k and 8k context. Reported gains reach roughly twofold over FlashAttention and 225 TFLOPs/s/GPU in training.
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
0/21This paper attributes delayed generalization in grokking partly to unequal optimization speeds along gradient singular directions. It introduces Egalitarian Gradient Descent (EGD), which replaces each selected layer's gradient by its polar factor so that nonzero singular values are equalized. A solvable anisotropic linear-classification model links covariance conditioning and initialization scale to plateau length, while experiments compare EGD, randomized-SVD approximations, column normalization, standard optimizers, and Grokfast. Across modular arithmetic, sparse parity, MNIST, and CIFAR-10 distribution-shift settings, the paper reports earlier generalization and improved adaptability, alongside computation, rank-sensitivity, instability, and trajectory analyses.
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
0/57Large language model agents frequently struggle with procedural coherence over long execution horizons, leading to disorganized tool use and redundant actions. This paper introduces the Procedural Graph (PG), an explicit, editable directed graph structuring task procedures into typed triplets annotated with execution conditions, guidance, and pitfalls. At inference, active nodes are localized to provide generative, step-level situational guidance from surrounding subgraphs. Offline, an LLM refiner evolves the graph topology and edge attributes via contrastive trajectory analysis, validation gating, and rejection memory. Evaluated across seven benchmarks, four frontier LLM families, and diverse domains, PG achieves consistent gains over memory baselines, constructs effective graphs from minimal skeletons, and repairs flawed expert priors.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
0/76Knowledge distillation with post-trained teachers is increasingly used to compress language models, yet whether its benefits remain consistent across training stages is unclear. Evaluating logit-based distillation on OLMo-2, the authors identify a stage-dependent reasoning-recall tradeoff: while forward-KL distillation improves both reasoning and factual recall during pre-training, it slows factual acquisition during mid-training. This discrepancy stems from asymmetric teacher confidence across domains—teachers are confident on procedural tasks but uncertain on factual tokens, which students learn early. To resolve this imbalance, the authors introduce Switch Distillation, dynamically routing confident tokens to reverse-KL distillation and uncertain tokens to cross-entropy. Switch Distillation matches or exceeds baselines across reasoning and knowledge benchmarks while preserving factual recall.
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
0/14Skill-augmented agents persist loaded skills across execution loops, creating severe vulnerabilities where malicious or compromised skills steer tool calls, leak secrets, or corrupt state after installation. Static pre-installation vetting cannot detect these task-conditioned threats. To address this challenge, the authors introduce Defense-as-Skill, deploying the runtime guard as an installable, inspectable, and editable skill named SkillSonar that dynamically gates sensitive actions through allow, replan, or confirmation decisions. They construct SCOPE-R, a benchmark spanning six risk families with 206 attack-confirmed malicious instances and 43 benign tasks, and optimize SkillSonar using feedback-driven Monte Carlo Tree Search. On repeated GLM-5 evaluations, SkillSonar substantially curtails attack success while preserving benign task utility across multiple agent harnesses.
Generalization Dynamics of LM Pre-training
0/21This study tracks language-model generalization during pretraining using six behavioral evaluations and additional fine-tuning tests. Across OLMo3 and Apertus checkpoints, models repeatedly switch between following shallow patterns and producing answers consistent with task structure, even late in training. The reported fluctuations persist with probability-based metrics and alternative training-compute axes, while common benchmarks remain comparatively stable. Single-step updates and checkpoint averaging do not eliminate the phenomenon. Larger models generalize more often, but still fluctuate. Intermediate checkpoint selection improves post-training reasoning transfer and resistance to prefilling attacks in the reported comparisons. A preliminary data-selection experiment also steers generalization dynamics. Complexity proxies give inconsistent predictions across layers and datasets.
Just-In-Time Agent Memory with Runtime Agentic Research
0/18This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.
Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
0/66This paper studies why coordinate-based multilayer perceptrons struggle to fit high-frequency signals in low-dimensional domains. It analyzes training through neural tangent kernels and shows how sinusoidal Fourier features produce a stationary effective kernel whose bandwidth can be adjusted. Experiments examine convergence, generalization, feature sampling distributions, network depth, joint feature optimization, translation sensitivity, and directional bias. Direct image and shape regression and indirectly supervised CT, MRI, and simplified NeRF reconstruction compare unembedded inputs with basic, positional, and Gaussian mappings. Gaussian features provide the strongest reported results among the main mappings, while bandwidth selection balances underfitting and overfitting. Appendix studies identify limitations of feature optimization and axis-aligned encodings.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
0/41The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.
Online Self-Weighted Fine-Tuning
0/5Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
LongPIBench: A Long-Context Benchmark for Prompt Injection
0/4LongPIBench introduces a benchmark for prompt-injection attacks and defenses in four document-centric long-context workflows: paper peer review, resume screening, email summarization, and code review. It combines synthetic and real-world datasets, evaluates heuristic and optimization-based attacks across eight language models, and measures both prevention-based attack success and detection-based false-positive and false-negative rates. The paper reports that attacks remain effective and that defenses that perform well on short-context benchmarks often degrade on long documents, with additional ablations over document format, injection position, attack goals, context length, and segmentation.
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
0/14LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.