Hi,
There are two more papers which I believe are super related :)
- KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments (https://arxiv.org/abs/2504.15364)
- CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction (https://arxiv.org/abs/2504.14051)
It would be great if you could consider adding them
Hi,
There are two more papers which I believe are super related :)
It would be great if you could consider adding them