Back to the map
Reinforcement Learning2026

Learned KV-Cache Eviction

A PPO agent that learns what to evict from a frozen LLM's cache, one decoding step at a time.

Overview

A final project for UdeSA's Reinforcement Learning course. It reformulates KV-cache eviction for a frozen LLM (Qwen2.5-1.5B-Instruct) as a sequence of discrete, maskable decisions, one per decoding step, instead of Apple's one-shot KVP ranking. Framed as a Gymnasium environment and trained with MaskablePPO, the policy chooses which cached tokens to evict at every step under a fixed memory budget, evaluated on GSM8K, HotpotQA, and a synthetic passkey-retrieval task.

Highlights

  • A thirteen-experiment ablation ladder (rich features, warm-start, cross-token attention, dense causal reward, per-layer credit) that all converge to parity with the kv_norm heuristic on GSM8K.
  • A future-attention oracle shows that parity is a property of the dataset, not the method: the learnable margin over the heuristic is +0.07 on GSM8K versus +0.43 on synthetic passkey retrieval and +0.09/+0.19 on HotpotQA, growing with compression aggressiveness.
  • In the signal-rich passkey arena, an online formulation with dense causal reward shows the first sustained edge over the heuristic, approaching the oracle eviction policy.

Report

Full report (PDF)

Read the report ↗
Next projectFastSLAM & Autonomous Navigation →