浏览 ArkGraph,选择本次要执行的步骤。
KV-COBRA studies extreme-rate compression of transformer key-value caches by treating rank and bit width as a coupled allocation problem. It selects a rank–precision pair for each attention head, redistributes total storage across heads, and uses a fused Hadamard rotation to equalize retained-channel variance. A query-aware variant reorders SVD directions using an attention-KL importance proxy. Experiments across three 7–8B models, additional 70B and mixture-of-experts models, perplexity, reasoning, long-context QA, and retrieval show the largest gains at low bit rates. Appendices examine solver behavior, quantizer-model error, calibration robustness, latency primitives, value-cache limits, RoPE placement, and depth sensitivity.
The paper argues that deciding where cache bits go can matter more than inventing a new quantizer, especially below one bit per dimension. In plain terms, different attention heads and directions tolerate compression very differently, and KV-COBRA exploits that variation. The evidence spans several model families and tasks, but it does not establish end-to-end latency, does not combine with token eviction or entropy coding, and depends on calibration-time approximations whose finite-bit errors and layer-depth sensitivity are documented.
论文中的结论可在「研究结论」中查看。