Skip to results
MLSift
← Feed
routineML Systems & EfficiencyKV-cache compression2607.09683

Ablation, Statistical Inference, and Validation for KV-Cache Compression

Paolo D'Alberto, Ashish Siarasao, Elliott Delaye, Rajeev Patwari

cs.LG cs.AI cs.IT

Abstract

This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through a statistical validation methodology that separates systematic codec differences from implementation variance. Key findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, with the effective semantic dimension ($d_{eff}$) adapting to calibration budgets rather than true data rank. (this is an abstract of the abstract thank you )

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF