Public research repositories on CiteArk, sorted by activity, update time, claims with supporting Assessments, and community reproduction requests.
Subject Distributed, Parallel, and Cluster Computing · 3 repositories
ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
0/23ASPIRE is a batched self-speculative decoding system for long-context language-model inference. It lets requests independently draft with sparse attention or verify with full attention inside one target-model forward pass, schedules each request using online acceptance and cost estimates, and refreshes sparse context during drafting through one full-attention layer. Experiments span three open models and five reasoning or long-context workloads. The paper reports 1.70–4.58× decode-throughput speedups over autoregressive decoding, generally stronger gains at longer contexts, small greedy-quality differences, and robustness studies covering oracle scheduling, tensor parallelism, refresh position, page size, estimator constants, and draft length.
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent–Tool Boundary
0/15This paper studies consistency failures that arise when agent workflows invoke independently supplied tools whose external effects may be uncertain, irreversible, concurrent, or visible before workflow resolution. It introduces an effect-history model that separates real-world events from runtime observations, organizes eight anomalies into uncertainty, workflow, and interaction families, and derives boundary capabilities and transactional contract families needed to exclude them. It also compares six agent systems against the catalog and surveys standard annotations for 98,291 tools from remotely reachable Model Context Protocol servers. The census finds widespread but coarse annotation use and concludes that the current vocabulary cannot fully express any required transactional capability.
LoGo: Token-Level Dynamic Local-Global Attention
0/6The paper introduces LoGo, a decoder-only Transformer attention mechanism that gives every token a local causal-attention branch and selectively activates a full-context branch through a learned token-level gate. An adaptive threshold controls the global activation ratio without an auxiliary balancing loss, progressive masking stabilizes training, and query-sparse Triton kernels avoid computing global attention for unselected queries. Experiments compare LoGo with full-attention and static local-global hybrids across model scales, long-context extension stages, recall benchmarks, kernel runtimes, ablations, and routing analyses.