Pinned Tweet
After almost 12 years in Brain/DeepMind, I’ve finally decided to take the leap. My cofounders: @yinfeiy, Seth and I have kicked-off @ElorianAI. The first multimodal reasoning lab founded and led by former LLM pretraining, data and multimodal leads. youtu.be/YlvfNpOMeOY?si=iiVT… (1/n)
What is visual reasoning? Just watch an engineer diagnose a failure. They connect dots across drawings, thermal images, simulations, and inspections to see what’s wrong and find a solution.
MechReason (arxiv.org/html/2609.16012v1) shows models fail most at this stage. At @ElorianAI, we’re building models that reason like engineers.
Recognizing a pattern is easy. Understanding the structure behind it is much harder.
Recent benchmarks show models can identify patterns in image sequences but constantly select the wrong image to complete them.
The models recognized visual similarities but failed to reason about underlying relationships well enough to predict what came next.
That’s the gap we’re working to close at @ElorianAI. We’re building models that can move beyond visual pattern matching toward genuine structural reasoning.
Reasoning first evolved from vision and visual understanding. That's an entire space of reasoning that we haven't yet explored.
Everyone knows that next-token prediction scales for language but you need to think outside to find the objective that scales for video and understanding the physical world.
Why didn't anyone ping us about pacing the frontier? Then we wouldn't have had to cook so hard this weekend.
Next token prediction doesn't work because it accurately models how human brains learn. It works because it's one of the only objectives that scale, unlike some other ones.
In the early days of OpenAI, @johnschulman2 didn't think next-token prediction was going to lead to intelligence, because it would get swamped by noise.
He explains it's always been hard to apriori predict what techniques will elicit out-of-distribution generalization.
OSU's Professor Hai-Jun Su has been visiting @ElorianAI this summer investigating the combination of mechanical engineering and VLM research — and it was clear he was onto something big. He's since been recognized and awarded at ASME IDETC-CIE 2026 three times for his work on new ways to design machines and robots, and honestly, no one deserves it more. Great work, Professor! mae.osu.edu/news/2026/09/mae…
Excited that we now have AI that achieves 3 year old AGI with ARC-AGI-3 (64x64)! Really looking forward to Atari-game AGI (160x192), ARC-AGI-4. AGI Chess grandmaster, ARC-AGI-13. And maybe eventually AGI Starcraft, tentatively ARC-AGI-16 (640x480)?
As AI evolves, a model’s ability to comprehend the contents of video and images will only increase. People might say things like counting aren't important but any teacher on a field trip would disagree. I recently had the pleasure of chatting with Ben Lorica about @ElorianAI's progress in building models that understand complexities in video and counting things.
Check out my episode of The Data Exchange podcast to hear more about what we’re working on:
youtube.com/watch?v=iu-sdY_x…
Last week I was at @AISummitSeoul on a panel with @jeffclune and @SeoulNatlUni. This week? Ray Summit. I shared the same message in both rooms, that visual AI still needs massive improvements. But we're still in the early days. Once we get it right, the upside for engineers, designers, architects, and more is huge.
Excited to present at the #RaySummit today about what we're working on at @ElorianAI !
At #RaySummit 2026—robotics track is more than 2025. 🤖
→ @nvidia Isaac Lab × @anyscalecompute @raydistributed: thousands of parallel humanoid loco/nav sims
→ @LeRobotHF v3.0: dataset layout becomes a scaling bottleneck
→ @mimicrobotics: Video Action Models—pretrained video backbones + human video for dexterous control
→ @AndrewDai: why frontier VLMs still can’t actually see
Physical AI is increasingly an infra problem: sim throughput × data layout × video priors × visual grounding.
#RaySummit #PhysicalAI
We're hiring at Elorian. We're solving a problem that's still wide open: teaching AI to think about the visual realm the way humans do. If you want to join a focused, highly selective team and help define what's next, we'd love for you to join us: elorian.ai
Google Meet translated captions are amazingly useful. And can be amazingly funny too... seen firsthand on a Korean business call.
I’ve been obsessed with board games since lockdown. I love the games themselves but I also love the pieces, how a plain colored wooden token can carry so much meaning or how the icons on a card can tell you everything you need to know with zero words.
I think about this a lot in our work and it’s this rule, that understanding the world can often come via the most basic cues, that inspired our logo, new brand identity and website.
Our new identity centers around the idea that visual art - in it's most natural form, free from low-res or grid-based constraints - often leaves room for you to bring your own interpretation. Elorian’s new logo and website show how with just the right amount of information, humans can fill in the rest.
I’m curious what it brings up for you: elorian.ai
I’m leaving Meta to start a new company.
Building the TBD Lab alongside Mark @finkd and Alex @alexandr_wang has been deeply inspiring and fulfilling. I’m proud of what our multimodal team accomplished across Muse Spark, Voice Mode, Muse Image, and Muse Video, and even prouder of the team that made it possible.
Over time, I’ve felt increasingly drawn to a problem that will matter deeply to humanity’s future, yet remains largely underexplored. It now has my full attention.
More to share as the work takes shape.
Another @JeffDean anecdote: in the pre-TensorFlow async training days, we were investigating why parameter updates were slow. A typical engineer would only read the C++ code, Jeff dug all the way into assembly code and found the extraneous memory load/lock instruction.
He also gave a great piece of advice: don't spend more than 3 or so years in a single project. Keep changing projects so you can keep learning. Great to see him stick to that!
Jeff built a curiosity-driven and open research culture at Brain unlike anywhere else. It's the direct inspiration to the culture we're building at @ElorianAI.