Back to the mapRead the report ↗
Interpretability & Reasoning2026
Adaptive Latent Reasoning
Letting a model decide how many latent reasoning steps a problem deserves.
Overview
Latent chain-of-thought methods reason in continuous space instead of emitting tokens, but they fix the number of reasoning steps in advance: the same budget for an easy sum and a hard word problem. This NLP final project extends SIM-CoT (ICLR 2026) with PonderNet-style adaptive halting so the number of implicit steps becomes variable at inference time, learned per input. The model spends more latent computation on harder GSM8K problems and stops early on easy ones.
Highlights
- Adds a learned halting distribution over latent reasoning steps on top of SIM-CoT / CODI, trained on GSM8K.
- Includes a from-scratch GPT-2 run for a fair comparison against fixed-step baselines (Coconut, CODI, SIM-CoT).
- Reproducible experiment pipeline with per-run logging and a documented experiment index.
Report
Full report (PDF)