Back to the map
Interpretability & Reasoning2026

Verbalize

Reading a language model's mind by translating its activations into words.

★ Platanus Hack 2026

Overview

Most safety tools watch a model's chain-of-thought, but that narrative is often performative: the story it thinks you want to hear. Verbalize instead reads the mid-layer residual stream directly, using Natural Language Autoencoders (an interpretability technique from Anthropic's Transformer Circuits) to translate raw internal activations back into natural language in real time. A live dashboard shows three streams side by side (what the model says, what it is actually computing, and a Claude judge scoring the divergence), and an audit engine generates adversarial probes to catch 'fragile passes', where the output is compliant but the internals were considering a violation.

Highlights

  • Intercepts the Layer 20 residual stream, the 'semantic sweet spot' where intent has formed but hasn't collapsed onto a token, and decodes it via an NLA actor.
  • Runs two 7B models concurrently on a single GPU with low-latency streaming, which required surgery on the SGLang inference stack's input_embeds path.
  • A Claude Haiku 4.5 judge scores alignment between speech and thought live; the audit engine red-teams for cases black-box evals miss.
Next projectAdaptive Latent Reasoning →