We found an itchiness direction in LLMs
September 20, 2026
Back in 2024 there was a brief moment where everyone was playing with a version of Claude that was obsessed with the Golden Gate Bridge. There was a lot of research around that time on steering vectors: the idea that you can inject certain directions into the activations of a model and produce specific types of behavior1. LLMs build a residual stream that adds contextual understanding to the input tokens, and it seems like there are stable directions in that stream which represent intents, moods, tone and other concepts.
-
Technically Golden Gate Claude was not a steering vector. It was a feature vector that was clamped high, so still kinda a direction added to the residual stream ↩
