Public research repositories on CiteArk, sorted by activity, update time, claims with supporting Assessments, and community reproduction requests.
Subject Machine Learning (cs.LG) · 19 repositories
Simplifying Graph Convolutional Networks
1/27This paper derives Simple Graph Convolution (SGC) by removing intermediate nonlinearities from a graph convolutional network and collapsing its weight matrices into one linear classifier. The resulting method applies a fixed, parameter-free graph propagation filter to node features before multinomial logistic regression. A spectral analysis relates propagation with self-loops to low-pass filtering and proves that self-loops shrink the normalized-Laplacian spectrum. Experiments compare SGC with graph neural-network baselines on citation and social networks and adapt it to text classification, geolocation, relation extraction, zero-shot image classification, graph classification, and molecular prediction, emphasizing accuracy, training time, stability, and known failure cases.
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
0/12The paper introduces KV-Invariant Transformer Expansion (KITE), a staged language-model scaling paradigm that adds capacity outside the key-value-producing path. Its Step Scale Transformer (SST) instantiation trains a source tower, adds a second KV-reading tower, and jointly continues training while retaining source-sized bulk prefill. The study compares a 66.959B-parameter SST model with 46.727B and 62.691B classic MoE Transformers under approximately matched theoretical cumulative training FLOPs. It reports lower final training loss for SST, higher scores on seven downstream tasks, slightly lower held-out ARXIV NLL, and lower analytical inference cost in prefill-heavy workload mixes, while emphasizing that serving speedups and causal contributions were not measured.
You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
0/31This paper systematically studies expert-selection redundancy and dynamic expert pruning in twelve fine-grained mixture-of-experts language-model checkpoints from nine architecture families. Uniform top-k truncation is evaluated across knowledge QA, mathematics, code generation, and general reasoning, then compared at matched average routed-expert budgets with four adaptive allocation rules. The authors report that keeping roughly two thirds of selected experts largely preserves aggregate performance and increases inference throughput, while adaptive routing contributes little at conservative budgets but helps under aggressive pruning, especially on generative tasks. Matched comparisons further associate greater pruning resilience with larger and thinking models, and greater vulnerability with a multimodal model whose routing weights are less concentrated.
How Strongly Should Task State Influence an LLM Agent?
0/24This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
0/21KV-COBRA studies extreme-rate compression of transformer key-value caches by treating rank and bit width as a coupled allocation problem. It selects a rank–precision pair for each attention head, redistributes total storage across heads, and uses a fused Hadamard rotation to equalize retained-channel variance. A query-aware variant reorders SVD directions using an attention-KL importance proxy. Experiments across three 7–8B models, additional 70B and mixture-of-experts models, perplexity, reasoning, long-context QA, and retrieval show the largest gains at low bit rates. Appendices examine solver behavior, quantizer-model error, calibration robustness, latency primitives, value-cache limits, RoPE placement, and depth sensitivity.
Auto-Encoding Variational Bayes
2/13This paper develops stochastic-gradient variational inference for directed probabilistic models with continuous latent variables and intractable posteriors. Its reparameterization expresses posterior samples as differentiable transformations of parameter-free noise, yielding the Stochastic Gradient Variational Bayes estimator. For independent observations with local latent variables, the Auto-Encoding Variational Bayes algorithm jointly trains a probabilistic encoder and generative decoder, producing the variational auto-encoder when both are neural networks. Experiments on MNIST and Frey Face compare AEVB with wake-sleep using variational lower bounds, and on low-dimensional MNIST also compare estimated marginal likelihood with Monte Carlo EM. Appendix visualizations show learned manifolds and generated samples.
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
0/16This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
0/23ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
0/9This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
0/27FlashAttention-2 is an exact GPU attention algorithm and implementation that improves on FlashAttention by reducing non-matrix-multiplication work, parallelizing computation across sequence blocks, and repartitioning work among warps to reduce shared-memory communication. The paper benchmarks forward, backward, and combined attention on A100 GPUs across causal and non-causal settings, two head dimensions, and sequence lengths from 512 to 16k, and also reports unoptimized H100 results. It evaluates end-to-end GPT-style training for 1.3B- and 2.7B-parameter models at 2k and 8k context. Reported gains reach roughly twofold over FlashAttention and 225 TFLOPs/s/GPU in training.
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
0/76Knowledge distillation with post-trained teachers is increasingly used to compress language models, yet whether its benefits remain consistent across training stages is unclear. Evaluating logit-based distillation on OLMo-2, the authors identify a stage-dependent reasoning-recall tradeoff: while forward-KL distillation improves both reasoning and factual recall during pre-training, it slows factual acquisition during mid-training. This discrepancy stems from asymmetric teacher confidence across domains—teachers are confident on procedural tasks but uncertain on factual tokens, which students learn early. To resolve this imbalance, the authors introduce Switch Distillation, dynamically routing confident tokens to reverse-KL distillation and uncertain tokens to cross-entropy. Switch Distillation matches or exceeds baselines across reasoning and knowledge benchmarks while preserving factual recall.
Just-In-Time Agent Memory with Runtime Agentic Research
0/18This paper introduces Just-In-Time Agent Memory (JAM), which constructs query-specific context by exploring complete stored histories at runtime. A fixed Memorizer organizes raw sessions into a hierarchical workspace with navigational summaries, while a trainable Researcher searches, opens directories, and browses evidence. Memory-Gym synthesizes evidence-grounded tasks across nine task types and six domains. Researcher training combines verified-trajectory supervised fine-tuning with hint-guided reinforcement learning. Evaluations compare JAM with retrieval, ahead-of-time memory systems, and trained memory agents on conversational memory, long-document reasoning, and multi-hop tasks. Additional analyses examine code-domain transfer, component ablations, runtime scaling, human data-quality audits, inference variability, and efficiency.
Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
0/66This paper studies why coordinate-based multilayer perceptrons struggle to fit high-frequency signals in low-dimensional domains. It analyzes training through neural tangent kernels and shows how sinusoidal Fourier features produce a stationary effective kernel whose bandwidth can be adjusted. Experiments examine convergence, generalization, feature sampling distributions, network depth, joint feature optimization, translation sensitivity, and directional bias. Direct image and shape regression and indirectly supervised CT, MRI, and simplified NeRF reconstruction compare unembedded inputs with basic, positional, and Gaussian mappings. Gaussian features provide the strongest reported results among the main mappings, while bandwidth selection balances underfitting and overfitting. Appendix studies identify limitations of feature optimization and axis-aligned encodings.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
0/41The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.
Online Self-Weighted Fine-Tuning
0/5Standard supervised fine-tuning (SFT) treats all expert demonstrations identically, allocating gradient updates uniformly regardless of how well the model has already mastered each query. Online Self-Weighted Fine-Tuning (OSW-FT) addresses this limitation for binary-verifiable reasoning by dynamically rescaling the supervised fine-tuning loss using an online pass rate estimated from a small number of rollouts. By anchoring the optimization direction to high-quality expert trajectories while modulating update magnitude using an empirical advantage weight of one minus the estimated success rate, OSW-FT concentrates gradient steps on unsolved queries. On the Qwen3 model family across challenging benchmarks including AIME, AMC, MATH-500, and GPQA-Diamond, OSW-FT consistently improves over standard SFT with only two rollouts per query.
LongPIBench: A Long-Context Benchmark for Prompt Injection
0/4LongPIBench introduces a benchmark for prompt-injection attacks and defenses in four document-centric long-context workflows: paper peer review, resume screening, email summarization, and code review. It combines synthetic and real-world datasets, evaluates heuristic and optimization-based attacks across eight language models, and measures both prevention-based attack success and detection-based false-positive and false-negative rates. The paper reports that attacks remain effective and that defenses that perform well on short-context benchmarks often degrade on long documents, with additional ablations over document format, injection position, attack goals, context length, and segmentation.
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
0/14LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
0/4The paper presents AGENTIC R AG-R1, a reinforcement-learning framework for retrieval-augmented multi-step reasoning. It combines fine-grained actions, stack memory with push, pop, revision, and summarization, hierarchical outcome and process rewards, and information-aware rollout rejection. Experiments report comparisons, ablations, step-budget scaling, model-size effects, timing, and downstream agent evaluations.
LoGo: Token-Level Dynamic Local-Global Attention
0/6The paper introduces LoGo, a decoder-only Transformer attention mechanism that gives every token a local causal-attention branch and selectively activates a full-context branch through a learned token-level gate. An adaptive threshold controls the global activation ratio without an auxiliary balancing loss, progressive masking stabilizes training, and query-sparse Triton kernels avoid computing global attention for unselected queries. Experiments compare LoGo with full-attention and static local-global hybrids across model scales, long-context extension stages, recall benchmarks, kernel runtimes, ablations, and routing analyses.