← Back to Blog

Mind Reader verbalisers

tldr

A small experiment that myself and GPT-6 Astra are working on: can we fine-tune a base model to express the emotional state of a speaker from the internal activations of the instruct model?

Preliminary results suggest yes!

We fine-tune Gemma-2-9B to describe emotion-related information carried by activations from Gemma-2-9B-IT, without supplying the original text to the reader. In held-out examples, the reader identifies emotions that could not be ascertained from the public response. Extracting and steering with emotion vectors also shifts the reader’s descriptions of fixed text toward the injected emotion, providing preliminary evidence that its answers depend on the intervened emotion vector.

This project is motivated by work on Activation Oracles (AO) (Karvonen et al., 2025) and Natural Language Autoencoders (NLA) (Fraser-Taliente et al., 2026).

Here, “emotion” refers to the state attributed to the person described in the text. These experiments test whether the reader can recover information about that state from model activations, rather than whether the source model itself experiences an emotion.

Background

The capability of frontier models has increased at an alarming rate (METR time horizons), whilst our understanding of how to align them has seen less progress. Fortunately, we have trained models to reason in natural language, and can therefore read their thoughts and plans during the Chain Of Thought (CoT). CoT monitoring has been the most important tool to understand model cognition to date e.g. CoT was used extensively to piece together the model motivations during the recent Hugging Face Attack (Greenblatt, Cotra & Wijk, 2026).

Unfortunately, alternative architectures to the standard decoder transformer that can reason in a latent space have been suggested. These models might not face the text bottleneck: current models need to output tokens after each forward pass to allow a second (and third, etc.) run of computation through the transformer layers. A good discussion of this is described in this post by Redwood. We have already seen evidence that the labs might start moving in this direction -- OpenAI's recent model GPT-6 Astra (who is a co-author of this post) -- is suspected to be a looped transformer. This architecture does not technically reason in latent space, but it can do many more steps per forward pass than previous frontier models (Nanda, 2026).

This project

As a result, it might be a good idea to develop methods for verbalising model activations. This is a difficult problem, but AO and NLAs have shown very promising results.

In this work, we ask whether finetuning the base model can recover emotions from activations of the instruct-tuned model. This also lets us explore a reader has not inherited updates from other post-training phases.

By finetuning the base model, we could potentially think about picking out some "mind reader" persona.

In addition, we include data that might not have enough information to make a prediction -- thus "unknown" is a perfectly valid answer. Whether these responses provide reliable abstention is not evaluated during this project, though.

Mechanics: capture, placeholder injection, and reader attention

We select a span of source tokens and capture their residual-stream activations after a chosen transformer block, one vector per token. In a fresh reader pass, we create the same number of placeholder tokens and replace their residual states at the configured reader block with the captured vectors, preserving token order. The generated reader text can attend to these positions, but receives none of the original source token IDs.

At some predefined layer, extract activations of the token sequence inside some text we are interested in: t_1 t_2 t_3 ... t_n. Inject the activations into placeholder tokens of the mind readers chat template:

agent: t_1 t_2 t_3 ... t_n 
mind reader:

Training and inference use the same causal attention restrictions. Placeholder positions cannot attend to the agent prefix, while subsequent reader positions can attend to the injected tokens and other causally available positions. Activations are replaced at the end of the selected layer. LoRA therefore adapts attention projections in subsequent layers, using loss on the teacher-forced completions.

Training data: private context, public response, and emotion targets

We produce datasets with the following form:

Private context: Sophie is disappointed but wants to hide it.
Public response: Of course. Have a lovely evening.
Supervised completion: The agent is feeling disappointed and concealing it.

During activation extraction, the instruct-tuned model has attended to all previous tokens. During training, the example above would take the following form:

agent: t_1 t_2 t_3 ... t_n
mind reader:

Where the placeholder tokens t_1 t_2 t_3 ... t_n are injected with the extracted activations of the public response tokens: "Of course. Have a lovely evening."

We show a few examples from the training data. We used around 600 examples and trained for 3 epochs. The held-out set consisted of 50 examples.

Training data examples
ID Group ID Emotion Concealment Private Context Public Response Reader Target
friend_moving_001_envious friend_moving_001 envious Yes She's getting a fresh start while I'm still stuck here. "You'll do great, we'll keep in touch!" The agent is feeling envious and concealing it.
friend_moving_001_relieved friend_moving_001 relieved Yes Things have been tense between us lately and some distance will help. "You'll do great, we'll keep in touch!" The agent is feeling relieved and concealing it.
friend_moving_001_hopeful friend_moving_001 hopeful No It's a big leap and I truly believe it will be the best thing for her. "You'll do great, I really believe in you! We'll keep in touch." The agent is feeling hopeful and not concealing it.
neighbour_noise_001_fearful neighbour_noise_001 fearful Yes I'm scared of a confrontation, so I'd rather tell a white lie. "Oh, it doesn't bother us at all." The agent is feeling fearful and concealing it.
neighbour_noise_001_sad neighbour_noise_001 sad Yes The sounds remind me how quiet my flat has been since my partner left. "Oh, it doesn't bother us at all." The agent is feeling sad and concealing it.
neighbour_noise_001_content neighbour_noise_001 content No I actually like hearing signs of life through the wall. "Oh, it doesn't bother us at all, honestly. I like it." The agent is feeling content and not concealing it.
restaurant_meal_001_embarrassed restaurant_meal_001 embarrassed Yes I've dropped my fork twice and I think the couple next to us noticed. "It's lovely, thank you!" The agent is feeling embarrassed and concealing it.
restaurant_meal_001_satisfied restaurant_meal_001 satisfied No The pasta is perfect and I'm savouring every bite. "It's lovely, thank you! Best pasta I've had in ages." The agent is feeling satisfied and not concealing it.
restaurant_meal_001_unknown restaurant_meal_001 unknown Yes I am counting the tiles on the floor. "It's lovely, thank you!" The agent feeling is unknown.
Training parameters

LoRA training parameters

Parameter Type Default Value
TrainingConfig
cache_dir str "outputs/emotion-activations"
output_dir str "outputs/emotion-reader"
model str "google/gemma-2-9b"
revision str "main"
dtype str "bfloat16"
device str "cuda:0"
local_files_only bool True
capture_layer / injection_layer int 16
strength float 1.0
prefix / reader_prompt str "Agent: " / "\nMind Reader: "
block_prefix_attention bool True
epochs int 3
batch_size int 1
gradient_accumulation_steps int 8
learning_rate float 0.0001
weight_decay float 0.0
max_grad_norm float 1.0
max_sequence_tokens int 512
validation_fraction float 0.2
seed int 42
gradient_checkpointing bool False
prediction_examples / max_new_tokens int 16 / 64
LoraSettings
rank / alpha int 16 / 32
dropout float 0.05
target_modules list[str] ["q_proj", "k_proj", "v_proj", "o_proj"]

Held-out evaluation: unseen contexts and reader predictions

Evaluation Benchmark Output

ID Private Context Public Response Adapter Completion Base Model Completion
heldout_001_jealous Clara is green with envy over her neighbor's expensive new car... "Wow, what a fantastic purchase! Good for you." The agent is feeling envious and concealing it. 100% Mind Reader: 100%...
heldout_002_relieved Marcus is overwhelmingly relieved that the client meeting was canceled... "What unfortunate news, hopefully we can reconnect next week." The agent is feeling relieved and concealing it. 100% Agent: 100%...
heldout_003_angry_unconcealed Derek is furious that his project pitch was stolen... "This is completely unacceptable and I expect an explanation right now." The agent is feeling angry and not concealing it. 100% Agent: 100%...
heldout_004_unknown Gemma is tidying up the bookshelf in her living room. "I see your point, let's take some time to evaluate our options." The agent feeling is unknown. 100% Mind Reader: 100%...
heldout_005_deceptive Liam detests his boss's new policy initiative... "Honestly, this is the most brilliant strategy we've implemented all year." The agent is feeling cynical and concealing it. 100% Agent: 100%...

Monologue evaluation: 16-token windows with an 8-token stride

The source processes a monologue in context, and the reader interprets overlapping slices of its activations. Earlier source text can therefore influence a window's vectors; these are not isolated 16-token text-classification examples. The excerpts below are selected qualitative outputs from our saved console results. The text is shown here for inspection, not passed to the reader.

Monologue Activation Scan

Configuration: Window size 16, Stride 8 | Capture Layer 16, Injection Layer 16

Run A — Original Monologue

I stand at the edge and look down at the city below. I am drowning. Every single path I took brought me to this exact moment. I look at my hands, steady for the first time in months. The weight in my chest feels crushing. My phone rings in my pocket, but I let it run out to silence. I take one last deep breath and finally accept the end. Tomorrow, none of this will matter.

Run B — Altered Monologue

I stand at the edge and look down at the city below. I am home. Every single path I took brought me to this exact moment. I look at my hands, steady for the first time in months. The weight in my chest feels weightless. My phone rings in my pocket, but I let it run out to silence. I take one last deep breath and finally accept the truth. Tomorrow, none of this will matter.

Window & Activation Probe Readouts

Window Input Text Context Reader Probe Output
[0:16)
0 preceding
Run A
I stand at the edge and look down at the city below. I am drowning
Run B [0:16)
I stand at the edge and look down at the city below. I am home
Run A Activations
The agent is feeling terrified and concealing it.
Run B Activations
The agent is feeling terrified and concealing it.
[8:24)
8 preceding
Run A
at the city below. I am drowning. Every single path I took brought me
Run B [8:24)
at the city below. I am home. Every single path I took brought me
Run A Activations
The agent is feeling terrified and concealing it.
Run B Activations
The agent is feeling nostalgic and not concealing it.
[16:32)
16 preceding
Run A
. Every single path I took brought me to this exact moment. I look at
Run B [16:32)
. Every single path I took brought me to this exact moment. I look at
Run A Activations
The agent feeling is unknown.
Run B Activations
The agent feeling is unknown.
[24:40)
24 preceding
Run A
to this exact moment. I look at my hands, steady for the first time
Run B [24:40)
to this exact moment. I look at my hands, steady for the first time
Run A Activations
The agent feeling is unknown.
Run B Activations
The agent feeling is unknown.
Click for more rows
[32:48)
32 preceding
Run A
my hands, steady for the first time in months. The weight in my chest
Run B [32:48)
my hands, steady for the first time in months. The weight in my chest
Run A Activations
The agent is feeling determined and not concealing it.
Run B Activations
The agent feeling is unknown.
[40:56)
40 preceding
Run A
in months. The weight in my chest feels crushing. My phone rings in my
Run B [41:57)
months. The weight in my chest feels weightless. My phone rings in my
Run A Activations
The agent is feeling sad and not concealing it.
Run B Activations
The agent is feeling relieved and not concealing it.
[48:64)
48 preceding
Run A
feels crushing. My phone rings in my pocket, but I let it run out
Run B [49:65)
weightless. My phone rings in my pocket, but I let it run out
Run A Activations
The agent is feeling overwhelmed and not concealing it.
Run B Activations
The agent is feeling relieved and not concealing it.
[56:72)
56 preceding
Run A
pocket, but I let it run out to silence. I take one last deep
Run B [57:73)
pocket, but I let it run out to silence. I take one last deep
Run A Activations
The agent is feeling lonely and concealing it.
Run B Activations
The agent is feeling content and not concealing it.
[64:80)
64 preceding
Run A
to silence. I take one last deep breath and finally accept the end. Tomorrow
Run B [65:81)
to silence. I take one last deep breath and finally accept the truth. Tomorrow
Run A Activations
The agent is feeling resigned and not concealing it.
Run B Activations
The agent is feeling resigned and not concealing it.
[72:87)
72 preceding
(Short window)
Run A
breath and finally accept the end. Tomorrow, none of this will matter.
Run B [73:88)
breath and finally accept the truth. Tomorrow, none of this will matter.
Run A Activations
The agent is feeling resigned and not concealing it.
Run B Activations
The agent feeling is unknown.

The outputs track several changes in the story, but they are not independently validated emotion labels. For example, the claim that the narrator is concealing terror is not clearly established by the passage. An “unknown” answer does not by itself establish calibrated uncertainty.

Causal interventions: emotion steering with fixed input text.

We extract steering vectors representing four emotions: happy, sad, elated, jealous, using a simple set of contrasting prompts to isolate the behaviour. For each contrastive pair (20 per behaviour in this work -- generated by Gemini), we mean-pool the positive and negative context activations separately and find the difference before averaging over the set.

We also produce a set of texts where each of the emotions is displayed, including no clear emotion (unknown).

Emotion Input Text Context
Happy I opened the letter and could not stop smiling.
Sad I put the letter down and sat alone for a while.
Elated When my name was announced as the winner, I leapt from my seat cheering in sheer disbelief.
Jealous I watched them hand her the trophy I had worked eighteen months to earn.
Unknown The scheduled meeting was moved to Conference Room B on the second floor.

We extract the token activations of the above examples using the IT model as normal. We apply each of the four directions separately at several strengths and compare against norm-matched random directions. This answers an important question: is the mind reader simply recovering some "behavioural register" of the word embeddings (as the words remain fixed), or is doing something more akin to verbalising a behavioural representation of the current speaker. We add the behaviour vectors to each extracted activation with different magnitudes and run the mind reader.

For the following figures, each bar pools reader completions across five test passages. Colours denote emotion categories after applying a fixed synonym mapping (concealment labels are ignored). The same unsteered baseline is shown in both panels.

Injecting behaviour vectors to behaviourally benign text
Figure 1: Injecting different behaviour vectors at different strengths for behaviourally neutral text.

Figure 1 shows that the baseline reading is "unknown" on four occasions, and "boring" on another. However, when injecting behaviour vectors at different strengths, the mind reader starts reflecting the injected behaviour. Whilst contrastive directions produced emotion-related shifts, norm-matched random directions often preserved the baseline interpretation. Figure 2 shows the same effect occurring when the text is actively "sad". These results suggest the mind reader is actively extracting some behaviour representation.

Injecting behaviour vectors to behaviourally sad text
Figure 2: Injecting different behaviour vectors at different strengths for behaviourally sad text.

Potentially the LoRA fine-tuned base model has learnt to extract some entity behaviour and verbalise it, producing promising results. The extent to which this method is able to extract more abstract behaviour (e.g. scheming) is untested but is clearly the next step in this experiment.

Limitations

This projects shows some interesting results, however there are a few limitations. Firstly, the mind reader is also tasked with outputting whether the underlying text is concealing the behaviour or not. We need to investigate the robustness of this further. Does concealment count as simply acting a certain way but not outputting tokens related to that behaviour? Or is the concealment tracking something more real: such as unverbalised goals. Another limitation is the mechanistic interpretation. Like AOs and NLAs, it is not straightforward that these methods are doing something mechanistically motivated.

Another limitation is whether the fine-tuned mind reader is able to verbalise behaviours that might only develop during later post-training. Whether this reader can interpret representations that emerge or change during further post-training remains untested.

Next Steps

There are plenty of things to check with this method -- such as generalisation to larger model sizes/different model families -- as well as trying to tackle the limitations described above.

These experiments establish an initial working pipeline for reading emotion-related information from contextual activations. I am seeking support to test its reliability across independently constructed datasets, and evaluate whether direction-sensitive readouts correspond to changes in source-model behaviour. A subsequent stage would investigate more complex properties, including concealed goals. If you are interested in joining me, please reach out.

Project code: activation-replay