Co-founder and CEO at Elorian AI. Ex (director) at Google DeepMind after 12y. Gemini Data Area Lead. 戴明博

Palo Alto, CA
Joined February 2011
After almost 12 years in Brain/DeepMind, I’ve finally decided to take the leap. My cofounders: @yinfeiy, Seth and I have kicked-off @ElorianAI. The first multimodal reasoning lab founded and led by former LLM pretraining, data and multimodal leads. youtu.be/YlvfNpOMeOY?si=iiVT… (1/n)
81
84
50
861
422,199
What is visual reasoning? Just watch an engineer diagnose a failure. They connect dots across drawings, thermal images, simulations, and inspections to see what’s wrong and find a solution. MechReason (arxiv.org/html/2609.16012v1) shows models fail most at this stage. At @ElorianAI, we’re building models that reason like engineers.
3
35
4,083
Recognizing a pattern is easy. Understanding the structure behind it is much harder. Recent benchmarks show models can identify patterns in image sequences but constantly select the wrong image to complete them. The models recognized visual similarities but failed to reason about underlying relationships well enough to predict what came next. That’s the gap we’re working to close at @ElorianAI. We’re building models that can move beyond visual pattern matching toward genuine structural reasoning.
1
27
1,443
Reasoning first evolved from vision and visual understanding. That's an entire space of reasoning that we haven't yet explored.
2
1
39
3,535
Everyone knows that next-token prediction scales for language but you need to think outside to find the objective that scales for video and understanding the physical world.
5
43
4,820
Why didn't anyone ping us about pacing the frontier? Then we wouldn't have had to cook so hard this weekend.
3
1
72
6,581
Next token prediction doesn't work because it accurately models how human brains learn. It works because it's one of the only objectives that scale, unlike some other ones.
In the early days of OpenAI, @johnschulman2 didn't think next-token prediction was going to lead to intelligence, because it would get swamped by noise. He explains it's always been hard to apriori predict what techniques will elicit out-of-distribution generalization.
2
3
1
110
21,215
Anyone who has trained frontier models can relate to this.
Wife told me these wouldn’t fit. Little did she know I had trained for this moment for years.
1
21
5,234
OSU's Professor Hai-Jun Su has been visiting @ElorianAI this summer investigating the combination of mechanical engineering and VLM research — and it was clear he was onto something big. He's since been recognized and awarded at ASME IDETC-CIE 2026 three times for his work on new ways to design machines and robots, and honestly, no one deserves it more. Great work, Professor! mae.osu.edu/news/2026/09/mae…
1
6
1,123
Excited that we now have AI that achieves 3 year old AGI with ARC-AGI-3 (64x64)! Really looking forward to Atari-game AGI (160x192), ARC-AGI-4. AGI Chess grandmaster, ARC-AGI-13. And maybe eventually AGI Starcraft, tentatively ARC-AGI-16 (640x480)?
3
22
3,679
As AI evolves, a model’s ability to comprehend the contents of video and images will only increase. People might say things like counting aren't important but any teacher on a field trip would disagree. I recently had the pleasure of chatting with Ben Lorica about @ElorianAI's progress in building models that understand complexities in video and counting things. Check out my episode of The Data Exchange podcast to hear more about what we’re working on: youtube.com/watch?v=iu-sdY_x…
1
11
4,209
More incredible work on math (the previous result led to a fields medal) by the folks at Axiom.
1/ ✨ World record today on bounded gaps between primes, one of number theory’s oldest open conjectures.  We show that infinitely many pairs of prime numbers are separated by 212 or less.
3
9
2,450
Last week I was at @AISummitSeoul on a panel with @jeffclune and @SeoulNatlUni. This week? Ray Summit. I shared the same message in both rooms, that visual AI still needs massive improvements. But we're still in the early days. Once we get it right, the upside for engineers, designers, architects, and more is huge.
5
33
2,142
Excited to present at the #RaySummit today about what we're working on at @ElorianAI !
At #RaySummit 2026—robotics track is more than 2025. 🤖 → @nvidia Isaac Lab × @anyscalecompute @raydistributed: thousands of parallel humanoid loco/nav sims → @LeRobotHF v3.0: dataset layout becomes a scaling bottleneck → @mimicrobotics: Video Action Models—pretrained video backbones + human video for dexterous control → @AndrewDai: why frontier VLMs still can’t actually see Physical AI is increasingly an infra problem: sim throughput × data layout × video priors × visual grounding. #RaySummit #PhysicalAI
1
3
21
3,345
We're hiring at Elorian. We're solving a problem that's still wide open: teaching AI to think about the visual realm the way humans do. If you want to join a focused, highly selective team and help define what's next, we'd love for you to join us: elorian.ai
64
28
7
549
56,740
Google Meet translated captions are amazingly useful. And can be amazingly funny too... seen firsthand on a Korean business call.
6
1,537
I’ve been obsessed with board games since lockdown. I love the games themselves but I also love the pieces, how a plain colored wooden token can carry so much meaning or how the icons on a card can tell you everything you need to know with zero words. I think about this a lot in our work and it’s this rule, that understanding the world can often come via the most basic cues, that inspired our logo, new brand identity and website. Our new identity centers around the idea that visual art - in it's most natural form, free from low-res or grid-based constraints - often leaves room for you to bring your own interpretation. Elorian’s new logo and website show how with just the right amount of information, humans can fill in the rest. I’m curious what it brings up for you: elorian.ai
3
36
3,471
Andrew M. Dai retweeted
I’m leaving Meta to start a new company. Building the TBD Lab alongside Mark @finkd and Alex @alexandr_wang has been deeply inspiring and fulfilling. I’m proud of what our multimodal team accomplished across Muse Spark, Voice Mode, Muse Image, and Muse Video, and even prouder of the team that made it possible. Over time, I’ve felt increasingly drawn to a problem that will matter deeply to humanity’s future, yet remains largely underexplored. It now has my full attention. More to share as the work takes shape.
294
123
57
4,006
703,176
Another @JeffDean anecdote: in the pre-TensorFlow async training days, we were investigating why parameter updates were slow. A typical engineer would only read the C++ code, Jeff dug all the way into assembly code and found the extraneous memory load/lock instruction. He also gave a great piece of advice: don't spend more than 3 or so years in a single project. Keep changing projects so you can keep learning. Great to see him stick to that!
27
84
1
2,147
394,811
Jeff built a curiosity-driven and open research culture at Brain unlike anywhere else. It's the direct inspiration to the culture we're building at @ElorianAI.
2
43
6,268