REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression
machine-learning compression retrieval transformers inference budget attention-mechanism reasoning eviction kv-cache attention-heads llm long-context cache-eviction longbench
-
Updated
Apr 15, 2026 - Python