
[Note] BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
aiha-lab/BeaconKV: BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference What problem? CoT (Chain-of-Thought) brings memory bottleneck (K...






