Back to the map
Interpretability & Reasoning2026

Adaptive Latent Reasoning

Letting a model decide how many latent reasoning steps a problem deserves.

Overview

Latent chain-of-thought methods reason in continuous space instead of emitting tokens, but they fix the number of reasoning steps in advance: the same budget for an easy sum and a hard word problem. This NLP final project extends SIM-CoT (ICLR 2026) with PonderNet-style adaptive halting so the number of implicit steps becomes variable at inference time, learned per input. The model spends more latent computation on harder GSM8K problems and stops early on easy ones.

Highlights

  • Adds a learned halting distribution over latent reasoning steps on top of SIM-CoT / CODI, trained on GSM8K.
  • Includes a from-scratch GPT-2 run for a fair comparison against fixed-step baselines (Coconut, CODI, SIM-CoT).
  • Reproducible experiment pipeline with per-run logging and a documented experiment index.

Report

Full report (PDF)

Read the report ↗
Next projectClaude on NIM →