Back to the mapRead the report ↗
Reinforcement Learning2026
Learned KV-Cache Eviction
A PPO agent that learns what to evict from a frozen LLM's cache, one decoding step at a time.
Overview
A final project for UdeSA's Reinforcement Learning course. It reformulates KV-cache eviction for a frozen LLM (Qwen2.5-1.5B-Instruct) as a sequence of discrete, maskable decisions, one per decoding step, instead of Apple's one-shot KVP ranking. Framed as a Gymnasium environment and trained with MaskablePPO, the policy chooses which cached tokens to evict at every step under a fixed memory budget, evaluated on GSM8K, HotpotQA, and a synthetic passkey-retrieval task.
Highlights
- A thirteen-experiment ablation ladder (rich features, warm-start, cross-token attention, dense causal reward, per-layer credit) that all converge to parity with the kv_norm heuristic on GSM8K.
- A future-attention oracle shows that parity is a property of the dataset, not the method: the learnable margin over the heuristic is +0.07 on GSM8K versus +0.43 on synthetic passkey retrieval and +0.09/+0.19 on HotpotQA, growing with compression aggressiveness.
- In the signal-rich passkey arena, an online formulation with dense causal reward shows the first sustained edge over the heuristic, approaching the oracle eviction policy.
Report
Full report (PDF)