Independent researcher studying mechanistic interpretability, internal state geometry, and dynamical systems

Joined November 2023
Super excited to announce the release of my new paper! Relational Rank Geometry in Transformers. Using Plücker coordinate geometry, I find that multi-token relations in Llama 3.1 8B, 70B, and 405B models leave higher order geometric frames in hidden-state space, and that steering those frames can recover clean answers in corrupt/clean tasks arxiv.org/abs/2605.29634
2
3
57
Also grateful to the researchers in mech interp and representation geometry who have been pioneering this way of studying model internals. It feels like such a powerful lens for understanding the structures that shape model behavior. @can_rager @EkdeepL @danielwurgaft @NeelNanda5 among many others
2
49
My sense is there many scientific breakthroughs that are waiting to happen, it's just that the subfields within science do not communicate effectively enough. My hope is that AI bridges that gap.
21
With every model release, I'm learning that sometimes getting out of the way is the best thing to when working with LLMs. I feel like I am becoming the bottleneck. Wild that this is feeling is only gonna get stronger the more we accelerate.
21
Took me a couple days to get codex for mobile for some reason. Now that I've got it up and running, idk if I should be excited for the productivity gains or concerned for my health now that I'll be running experiments all day. Oh well I guess.
29
Once we start to think of transformer's hidden states geometrically, we come closer to understanding their true nature.
Replying to @GoodfireAI
The same calculator handles a wide range of tasks, including: - arithmetic (“7+9”) - weekdays (“nine days after Friday”) - months (“six months after August”) Llama built this mechanism from scratch in training, and uses it with striking elegance and flexibility. (4/6)
23
Singular value decomposition might be one of the greatest representational tools for any dynamical system we have today.
13
Mazen Kobrosly retweeted
New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.
582
1,667
615
16,500
2,598,680
Mazen Kobrosly retweeted
Neural networks might speak English, but they think in shapes. Understanding their rich *neural geometry* is key to understanding how they work – and to debugging and controlling them with precision. Starting today, we’re releasing a series of posts on this research agenda. 🧵
310
1,674
513
11,312
3,431,530
Mazen Kobrosly retweeted
Can't explain it, but I trust GPT-5.5 more than Opus 4.7 right now.
345
78
45
3,707
765,994
Just came here to say codex is the shit. As someone who knows nothing about coding, it feels like absolute magic being able to build things out of thin air.
12
Trying to do this twitter thing and get my ideas out there. And collab with awesome people. Let's see how this goes!
11