flâneur — a map of the web's best reading

How do we know if activation verbalizers are telling us anything about activations? — Millicent Li

millicentli.github.io · 4,419 words · saved by 1 readers

TL;DR: Evaluations for activation verbalizers should establish simple baselines to inform us whether these verbalizers can actually yield information about the target LM, beyond what is accessible with black-box baselines. We find that evaluations for recent activation verbalizers, like Anthropic’s Natural Language Autoencoder (NLA), do not sufficiently meet this bar. Recently, there has been excitement around activation verbalization[1]: decoding the activations (or intermediate hidden states) of a target language model (LM) into natural language with another verbalizer LM (or similar model)[3, 4, 5, 6, 7, 8, 9, 10, 11]. The great hope of activation verbalization is to access a LM's "inner thoughts": Thoughts that LMs might never articulate outright. For instance, these thoughts should be knowledge available to the target LM but not directly accessible to the verbalizer without access to the target LM's internals. Activation verbalizers (also known under different names, such as Activ

How do we know if activation verbalizers are telling us anything about activations? — Millicent Li ← back to blog How do we know if activation verbalizers are telling us anything about activations? July 1st, 2026 TL;DR: Evaluations for activation verbalizers should establish simple baselines to inform us whether these verbalizers can actually yield information about the target LM, beyond what is accessible with black-box baselines. We find that evaluations for recent activation verbalizers, like Anthropic’s Natural Language Autoencoder (NLA), do not sufficiently meet this bar. Figure 1: I

Explore this link on the map →

related reading