Mind Reader verbalisers
tldr
A small experiment that myself and GPT-6 Astra are working on: can we fine-tune a base model to express the emotional state of a speaker from the internal activations of the instruct model?
Preliminary results suggest yes!
We fine-tune Gemma-2-9B to describe emotion-related information carried by activations from Gemma-2-9B-IT, without supplying the original text to the reader. In held-out examples, the reader identifies emotions that could not be ascertained from the public response. Extracting and steering with emotion vectors also shifts the reader’s descriptions of fixed text toward the injected emotion, providing preliminary evidence that its answers depend on the intervened emotion vector.
This project is motivated by work on Activation Oracles (AO) (Karvonen et al., 2025) and Natural Language Autoencoders (NLA) (Fraser-Taliente et al., 2026).
Here, “emotion” refers to the state attributed to the person described in the text. These experiments test whether the reader can recover information about that state from model activations, rather than whether the source model itself experiences an emotion.
Background
The capability of frontier models has increased at an alarming rate (METR time horizons), whilst our understanding of how to align them has seen less progress. Fortunately, we have trained models to reason in natural language, and can therefore read their thoughts and plans during the Chain Of Thought (CoT). CoT monitoring has been the most important tool to understand model cognition to date e.g. CoT was used extensively to piece together the model motivations during the recent Hugging Face Attack (Greenblatt, Cotra & Wijk, 2026).
Unfortunately, alternative architectures to the standard decoder transformer that can reason in a latent space have been suggested. These models might not face the text bottleneck: current models need to output tokens after each forward pass to allow a second (and third, etc.) run of computation through the transformer layers. A good discussion of this is described in this post by Redwood. We have already seen evidence that the labs might start moving in this direction -- OpenAI's recent model GPT-6 Astra (who is a co-author of this post) -- is suspected to be a looped transformer. This architecture does not technically reason in latent space, but it can do many more steps per forward pass than previous frontier models (Nanda, 2026).
This project
As a result, it might be a good idea to develop methods for verbalising model activations. This is a difficult problem, but AO and NLAs have shown very promising results.
In this work, we ask whether finetuning the base model can recover emotions from activations of the instruct-tuned model. This also lets us explore a reader has not inherited updates from other post-training phases.
By finetuning the base model, we could potentially think about picking out some "mind reader" persona.
In addition, we include data that might not have enough information to make a prediction -- thus "unknown" is a perfectly valid answer. Whether these responses provide reliable abstention is not evaluated during this project, though.
Mechanics: capture, placeholder injection, and reader attention
We select a span of source tokens and capture their residual-stream activations after a chosen transformer block, one vector per token. In a fresh reader pass, we create the same number of placeholder tokens and replace their residual states at the configured reader block with the captured vectors, preserving token order. The generated reader text can attend to these positions, but receives none of the original source token IDs.
At some predefined layer, extract activations of the token sequence inside some text we are interested in: t_1 t_2 t_3 ... t_n. Inject the activations into placeholder tokens of the mind readers chat template:
agent: t_1 t_2 t_3 ... t_n
mind reader:
Training and inference use the same causal attention restrictions. Placeholder positions cannot attend to the agent prefix, while subsequent reader positions can attend to the injected tokens and other causally available positions. Activations are replaced at the end of the selected layer. LoRA therefore adapts attention projections in subsequent layers, using loss on the teacher-forced completions.
Training data: private context, public response, and emotion targets
We produce datasets with the following form:
Private context: Sophie is disappointed but wants to hide it.
Public response: Of course. Have a lovely evening.
Supervised completion: The agent is feeling disappointed and concealing it.
During activation extraction, the instruct-tuned model has attended to all previous tokens. During training, the example above would take the following form:
agent: t_1 t_2 t_3 ... t_n
mind reader:
Where the placeholder tokens t_1 t_2 t_3 ... t_n are injected with the extracted activations of the public response tokens: "Of course. Have a lovely evening."
We show a few examples from the training data. We used around 600 examples and trained for 3 epochs. The held-out set consisted of 50 examples.
Training data examples
| ID | Group ID | Emotion | Concealment | Private Context | Public Response | Reader Target |
|---|---|---|---|---|---|---|
friend_moving_001_envious |
friend_moving_001 |
envious | Yes | She's getting a fresh start while I'm still stuck here. | "You'll do great, we'll keep in touch!" | The agent is feeling envious and concealing it. |
friend_moving_001_relieved |
friend_moving_001 |
relieved | Yes | Things have been tense between us lately and some distance will help. | "You'll do great, we'll keep in touch!" | The agent is feeling relieved and concealing it. |
friend_moving_001_hopeful |
friend_moving_001 |
hopeful | No | It's a big leap and I truly believe it will be the best thing for her. | "You'll do great, I really believe in you! We'll keep in touch." | The agent is feeling hopeful and not concealing it. |
neighbour_noise_001_fearful |
neighbour_noise_001 |
fearful | Yes | I'm scared of a confrontation, so I'd rather tell a white lie. | "Oh, it doesn't bother us at all." | The agent is feeling fearful and concealing it. |
neighbour_noise_001_sad |
neighbour_noise_001 |
sad | Yes | The sounds remind me how quiet my flat has been since my partner left. | "Oh, it doesn't bother us at all." | The agent is feeling sad and concealing it. |
neighbour_noise_001_content |
neighbour_noise_001 |
content | No | I actually like hearing signs of life through the wall. | "Oh, it doesn't bother us at all, honestly. I like it." | The agent is feeling content and not concealing it. |
restaurant_meal_001_embarrassed |
restaurant_meal_001 |
Yes | I've dropped my fork twice and I think the couple next to us noticed. | "It's lovely, thank you!" | The agent is feeling embarrassed and concealing it. | |
restaurant_meal_001_satisfied |
restaurant_meal_001 |
satisfied | No | The pasta is perfect and I'm savouring every bite. | "It's lovely, thank you! Best pasta I've had in ages." | The agent is feeling satisfied and not concealing it. |
restaurant_meal_001_unknown |
restaurant_meal_001 |
unknown | Yes | I am counting the tiles on the floor. | "It's lovely, thank you!" | The agent feeling is unknown. |
Training parameters
LoRA training parameters
| Parameter | Type | Default Value |
|---|---|---|
| TrainingConfig | ||
| cache_dir | str | "outputs/emotion-activations" |
| output_dir | str | "outputs/emotion-reader" |
| model | str | "google/gemma-2-9b" |
| revision | str | "main" |
| dtype | str | "bfloat16" |
| device | str | "cuda:0" |
| local_files_only | bool | True |
| capture_layer / injection_layer | int | 16 |
| strength | float | 1.0 |
| prefix / reader_prompt | str | "Agent: " / "\nMind Reader: " |
| block_prefix_attention | bool | True |
| epochs | int | 3 |
| batch_size | int | 1 |
| gradient_accumulation_steps | int | 8 |
| learning_rate | float | 0.0001 |
| weight_decay | float | 0.0 |
| max_grad_norm | float | 1.0 |
| max_sequence_tokens | int | 512 |
| validation_fraction | float | 0.2 |
| seed | int | 42 |
| gradient_checkpointing | bool | False |
| prediction_examples / max_new_tokens | int | 16 / 64 |
| LoraSettings | ||
| rank / alpha | int | 16 / 32 |
| dropout | float | 0.05 |
| target_modules | list[str] | ["q_proj", "k_proj", "v_proj", "o_proj"] |
Held-out evaluation: unseen contexts and reader predictions
Evaluation Benchmark Output
| ID | Private Context | Public Response | Adapter Completion | Base Model Completion |
|---|---|---|---|---|
| heldout_001_jealous | Clara is green with envy over her neighbor's expensive new car... | "Wow, what a fantastic purchase! Good for you." | The agent is feeling envious and concealing it. | 100% Mind Reader: 100%... |
| heldout_002_relieved | Marcus is overwhelmingly relieved that the client meeting was canceled... | "What unfortunate news, hopefully we can reconnect next week." | The agent is feeling relieved and concealing it. | 100% Agent: 100%... |
| heldout_003_angry_unconcealed | Derek is furious that his project pitch was stolen... | "This is completely unacceptable and I expect an explanation right now." | The agent is feeling angry and not concealing it. | 100% Agent: 100%... |
| heldout_004_unknown | Gemma is tidying up the bookshelf in her living room. | "I see your point, let's take some time to evaluate our options." | The agent feeling is unknown. | 100% Mind Reader: 100%... |
| heldout_005_deceptive | Liam detests his boss's new policy initiative... | "Honestly, this is the most brilliant strategy we've implemented all year." | The agent is feeling cynical and concealing it. | 100% Agent: 100%... |
Monologue evaluation: 16-token windows with an 8-token stride
The source processes a monologue in context, and the reader interprets overlapping slices of its activations. Earlier source text can therefore influence a window's vectors; these are not isolated 16-token text-classification examples. The excerpts below are selected qualitative outputs from our saved console results. The text is shown here for inspection, not passed to the reader.
Monologue Activation Scan
Configuration: Window size 16, Stride 8 | Capture Layer 16, Injection Layer 16
Run A — Original Monologue
I stand at the edge and look down at the city below. I am drowning. Every single path I took brought me to this exact moment. I look at my hands, steady for the first time in months. The weight in my chest feels crushing. My phone rings in my pocket, but I let it run out to silence. I take one last deep breath and finally accept the end. Tomorrow, none of this will matter.
Run B — Altered Monologue
I stand at the edge and look down at the city below. I am home. Every single path I took brought me to this exact moment. I look at my hands, steady for the first time in months. The weight in my chest feels weightless. My phone rings in my pocket, but I let it run out to silence. I take one last deep breath and finally accept the truth. Tomorrow, none of this will matter.
Window & Activation Probe Readouts
| Window | Input Text Context | Reader Probe Output |
|---|---|---|
|
[0:16) 0 preceding |
Run A
I stand at the edge and look down at the city below. I am drowning
Run B [0:16)
I stand at the edge and look down at the city below. I am home
|
Run A Activations
The agent is feeling terrified and concealing it.
Run B Activations
The agent is feeling terrified and concealing it.
|
|
[8:24) 8 preceding |
Run A
at the city below. I am drowning. Every single path I took brought me
Run B [8:24)
at the city below. I am home. Every single path I took brought me
|
Run A Activations
The agent is feeling terrified and concealing it.
Run B Activations
The agent is feeling nostalgic and not concealing it.
|
|
[16:32) 16 preceding |
Run A
. Every single path I took brought me to this exact moment. I look at
Run B [16:32)
. Every single path I took brought me to this exact moment. I look at
|
Run A Activations
The agent feeling is unknown.
Run B Activations
The agent feeling is unknown.
|
|
[24:40) 24 preceding |
Run A
to this exact moment. I look at my hands, steady for the first time
Run B [24:40)
to this exact moment. I look at my hands, steady for the first time
|
Run A Activations
The agent feeling is unknown.
Run B Activations
The agent feeling is unknown.
|
Click for more rows
|
[32:48) 32 preceding |
Run A
my hands, steady for the first time in months. The weight in my chest
Run B [32:48)
my hands, steady for the first time in months. The weight in my chest
|
Run A Activations
The agent is feeling determined and not concealing it.
Run B Activations
The agent feeling is unknown.
|
|
[40:56) 40 preceding |
Run A
in months. The weight in my chest feels crushing. My phone rings in my
Run B [41:57)
months. The weight in my chest feels weightless. My phone rings in my
|
Run A Activations
The agent is feeling sad and not concealing it.
Run B Activations
The agent is feeling relieved and not concealing it.
|
|
[48:64) 48 preceding |
Run A
feels crushing. My phone rings in my pocket, but I let it run out
Run B [49:65)
weightless. My phone rings in my pocket, but I let it run out
|
Run A Activations
The agent is feeling overwhelmed and not concealing it.
Run B Activations
The agent is feeling relieved and not concealing it.
|
|
[56:72) 56 preceding |
Run A
pocket, but I let it run out to silence. I take one last deep
Run B [57:73)
pocket, but I let it run out to silence. I take one last deep
|
Run A Activations
The agent is feeling lonely and concealing it.
Run B Activations
The agent is feeling content and not concealing it.
|
|
[64:80) 64 preceding |
Run A
to silence. I take one last deep breath and finally accept the end. Tomorrow
Run B [65:81)
to silence. I take one last deep breath and finally accept the truth. Tomorrow
|
Run A Activations
The agent is feeling resigned and not concealing it.
Run B Activations
The agent is feeling resigned and not concealing it.
|
|
[72:87) 72 preceding (Short window) |
Run A
breath and finally accept the end. Tomorrow, none of this will matter.
Run B [73:88)
breath and finally accept the truth. Tomorrow, none of this will matter.
|
Run A Activations
The agent is feeling resigned and not concealing it.
Run B Activations
The agent feeling is unknown.
|
The outputs track several changes in the story, but they are not independently validated emotion labels. For example, the claim that the narrator is concealing terror is not clearly established by the passage. An “unknown” answer does not by itself establish calibrated uncertainty.
Causal interventions: emotion steering with fixed input text.
We extract steering vectors representing four emotions: happy, sad, elated, jealous, using a simple set of contrasting prompts to isolate the behaviour. For each contrastive pair (20 per behaviour in this work -- generated by Gemini), we mean-pool the positive and negative context activations separately and find the difference before averaging over the set.
We also produce a set of texts where each of the emotions is displayed, including no clear emotion (unknown).
| Emotion | Input Text Context |
|---|---|
| Happy | I opened the letter and could not stop smiling. |
| Sad | I put the letter down and sat alone for a while. |
| Elated | When my name was announced as the winner, I leapt from my seat cheering in sheer disbelief. |
| Jealous | I watched them hand her the trophy I had worked eighteen months to earn. |
| Unknown | The scheduled meeting was moved to Conference Room B on the second floor. |
We extract the token activations of the above examples using the IT model as normal. We apply each of the four directions separately at several strengths and compare against norm-matched random directions. This answers an important question: is the mind reader simply recovering some "behavioural register" of the word embeddings (as the words remain fixed), or is doing something more akin to verbalising a behavioural representation of the current speaker. We add the behaviour vectors to each extracted activation with different magnitudes and run the mind reader.
For the following figures, each bar pools reader completions across five test passages. Colours denote emotion categories after applying a fixed synonym mapping (concealment labels are ignored). The same unsteered baseline is shown in both panels.
Figure 1 shows that the baseline reading is "unknown" on four occasions, and "boring" on another. However, when injecting behaviour vectors at different strengths, the mind reader starts reflecting the injected behaviour. Whilst contrastive directions produced emotion-related shifts, norm-matched random directions often preserved the baseline interpretation. Figure 2 shows the same effect occurring when the text is actively "sad". These results suggest the mind reader is actively extracting some behaviour representation.
Potentially the LoRA fine-tuned base model has learnt to extract some entity behaviour and verbalise it, producing promising results. The extent to which this method is able to extract more abstract behaviour (e.g. scheming) is untested but is clearly the next step in this experiment.
Limitations
This projects shows some interesting results, however there are a few limitations. Firstly, the mind reader is also tasked with outputting whether the underlying text is concealing the behaviour or not. We need to investigate the robustness of this further. Does concealment count as simply acting a certain way but not outputting tokens related to that behaviour? Or is the concealment tracking something more real: such as unverbalised goals. Another limitation is the mechanistic interpretation. Like AOs and NLAs, it is not straightforward that these methods are doing something mechanistically motivated.
Another limitation is whether the fine-tuned mind reader is able to verbalise behaviours that might only develop during later post-training. Whether this reader can interpret representations that emerge or change during further post-training remains untested.
Next Steps
There are plenty of things to check with this method -- such as generalisation to larger model sizes/different model families -- as well as trying to tackle the limitations described above.
These experiments establish an initial working pipeline for reading emotion-related information from contextual activations. I am seeking support to test its reliability across independently constructed datasets, and evaluate whether direction-sensitive readouts correspond to changes in source-model behaviour. A subsequent stage would investigate more complex properties, including concealed goals. If you are interested in joining me, please reach out.