JEPA enjoyer @BrownUniversity, prev building world models @overworld_ai, data @Bloomberg 🇱🇧/🇺🇸
Providence, RI
Joined March 2014
- Tweets263
- Following565
- Followers436
- Likes9.2K
Sami retweeted
Can a distributional loss pretrain a powerful one-step generative model from scratch?
The answer is YES but ONLY in the right feature space!
With the wrong representation (eg VAE features), training fails miserably and we uncover why in this work accepted to Neurips2026!
🧵
Sami retweeted
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.
But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.
I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.
I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.
But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
Sami retweeted
If you missed our opening remarks:
- 4th world modeling workshop: Aspen, CO, February, world models for physics, wmw-aspen.github.io
- 1st world modeling conference: Bay Area, CA, May, icwm.cc
In collaboration with @ylecun @LambdaAPI @amilabs and more TBA!
Sami retweeted
SIGReg for pretraining Video Foundation Models! Our LeVJEPA opens many doors...
- stable recipe with a simple loss (sigreg + prediction)
- no tubelet, frame aggregation, EMA, stop-gradient, ....
- 20X more FLOP efficient than VJEPA1/2 pretraining
- open source + reproducible
A new Pareto frontier in video pretraining.
Excited to introduce LeVJEPA 🔥: a stable, efficient end-to-end pretraining method that matches V-JEPA 2 at up to 20x less pretraining compute!
No target encoder, no masked prediction, no stop-gradient or teacher-student schedule.
One encoder, trained with a single objective. 🧵
Sami retweeted
Replying to @JitendraMalikCV
Exactly.
We also should not confuse world models, as per your definition, with video prediction or video generation models.
Understanding the dynamics of a system in order to control it is not the same as producing cute videos.
Sami retweeted
We are two weeks away from the third World Modeling Workshop!
- check wm-booth.org/ for all the latest info/links
- livestream and recordings will be available for free to everyone, no registration required
- a few surprises will be announced during the opening remarks
Sami retweeted
Action chunking is a mysteriously effective method. Modern large-scale imitation learning basically doesn't work without it. But why does it actually help? In our new paper, we try to break down the reasons. As the saying goes, what happened next might surprise you...
Action chunking is a critical component in virtually all modern approaches to imitation learning for robotics.
But why is it so critical, and do we really need action chunking? Check out our latest work to find out! (1/n)
action-chunking.github.io
Sami retweeted
We made an Elden Ring boss fight playable inside a real-time video world model.
Text is not a scene prompt. It is the controller.
Play Incantation:
reactor.inc/incantation
Sami retweeted
Today, I will give a talk about my recent Latent Action Model paper, including one possible recipe for for training Latent Action Models to learn Super Mario 🍄🕹️ from observation alone. Recordings will be available later on the same platform 🎥
Heejeong (Hazel) Nam of @BrownUniversity presents a novel approach that learns reusable motion primitives for more robust action understanding in AI.
🎙️ Hosted by @ceciletamura, Head of Community, @PloutosDev
📅 July 27, 2026
🕓 4:00 PM PDT
🔗 ploutos.dev/?stream=calm-ind…
Sami retweeted
Our humble recommendation- there is at least one better approach than EMA for latent space training arxiv.org/abs/2509.24317
why does JEPA/DINO-style training produce object-level structure in attention maps with zero segmentation supervision? one specific mechanistic hypothesis keeps coming up: the EMA teacher itself is doing a lot of the actual work
DINO's own ablation (Caron et al, arxiv.org/abs/2104.14294) is the cleanest evidence for this: swap the momentum-updated teacher for one that just copies the student's weights directly, and training fails to converge, while if you keeping the EMA, object boundaries emerge in the attention maps on their own, no segmentation labels anywhere in the loop
I-JEPA and the rest of the JEPA family inherited the same EMA teacher design for the same reason
my read on why this might matter mechanistically: a teacher that changes slower than the student acts as a moving but stable target, so the student can't just chase whatever the teacher did last step, it has to find features that stay predictive across many steps of teacher drift (objects are exactly the kind of feature that stays stable under small transformations while background noise doesn't, so a slow-enough target might be implicitly selecting for object-level structure rather than being told what one is)
still an open question how much of this is EMA specifically versus any sufficiently slow, sufficiently stable target (haven't found a paper yet that isolates the two cleanly)
Today's leading SSL methods rely heavily on training heuristics—EMA, teacher-student training, layer freezing, and more—to remain stable.
Regularization-based methods are much simpler, but have long failed to match the generalization performance of leading SSL approaches.
VISReg breaks this trade-off.
We introduce VISReg, the first regularization-based SSL method to achieve state-of-the-art out-of-domain generalization, outperforming MoCov3, DINO, data2vec, iBOT, I-JEPA, MAE, and DINOv2.
arXiv: arxiv.org/pdf/2606.02572
Github: github.com/HaiyuWu/visreg
Sami retweeted
Oops, SIGReg did it again! Large scale (CC12M->Datacomp-L) vision-language JEPA pretraining beats CLIP and SigLIP objectives! Thanks to SIGReg, our LeVLJEPA has no collapse, no EMA, no stop-gradient, no negatives, no problem! Checkpoints/demo are live: levljepa.github.io
🔥 We introduce LeVLJEPA: the first fully non-contrastive end-to-end vision-language pretraining method competitive with CLIP & SigLIP 💪🏼
👀 No negatives. No temperature. No momentum encoder. No teacher-student.
TL;DR: LeVLJEPA learns image to text structure by prediction: each modality predicts the other's embedding, while SIGReg keeps each embedding isotropic Gaussian. 🧵
📄 arxiv.org/abs/2607.00784
Sami retweeted
🔭 𝗘𝘅𝗽𝗹𝗼𝗿𝗮𝘁𝗶𝗼𝗻 in the Era of 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝘃𝗲 𝗖𝗼𝗻𝘁𝗿𝗼𝗹 🤖
Interacting with the world can be expensive!
Our #ICML2026 work shows how diffusion policies can 𝙚𝙭𝙥𝙡𝙤𝙧𝙚 during online experience collection to achieve sample-efficient self-improvement! 📈
Sami retweeted
Can regularization based JEPA (e.g. SIGReg) scale and compete with SOTA foundation models (DINO)? Here is the answer: yes and with 10x less data.
VISReg (slight variation of SIGReg) competes with DINOv2-LVD142M while only training on inet22k.
Try it out: huggingface.co/BooBooWu/visr…
Working on world model or SSL? You definitely need to try our new work: VISReg!
What does it achieve?
💪 Strong collapse prevention: High gradient when embedding collapse
⚡ Friendly to scale training: Linear complexity to scaling factors
🧩 Easy to train: Similar to LeJEPA, it is a heuristic-free method
🏆 Best OOD performance: Achieving the best accuracy on 6 OOD datasets
📉 Data efficiency: Achieving a similar OOD average accuracy to DINOv2 with 90% less data
🧬 Robust to low-quality datasets: It is robust to long-tailed and sparse datasets
Our results also indicate that SIGReg type methods can scale up, filling in the missing piece in @ylecun's great talk youtube.com/watch?v=72Xj8k5W….
A big thanks to my co-author @randall_balestr and my manager @DrMorganLevine. Also, huge gratitude to @ylecun for connecting us to make this project happen! 🤝
#SelfSupervisedLearning #JEPA #WorldModel
Sami retweeted
Can a Latent Action Model really know what the “action” is?
In Super Mario, Mario 👨🦰 moves but so do 🍄, ☁️,🌳, and the view🎥.
TL;DR: under agent ambiguity, don’t force LAMs to find the true action directly. We factorize what changed.
Paper: arxiv.org/abs/2606.30544v1
Sami retweeted
Would you like to join the research effort on JEPA and World Models easily?
After a full year of hard work, we’re excited to finally release stable-worldmodel:
an open-source, scalable platform built to accelerate JEPA & World Model research!
📄: github.com/galilai-group/sta…
another banger from the goat of goats
JEPA are finally easy to train end-to-end without any tricks!
Excited to introduce LeWorldModel: a stable, end-to-end JEPA that learns world models directly from pixels, no heuristics.
15M params, 1 GPU, and full planning <1 second.
📑: le-wm.github.io
Sami retweeted
JEPA are finally easy to train end-to-end without any tricks!
Excited to introduce LeWorldModel: a stable, end-to-end JEPA that learns world models directly from pixels, no heuristics.
15M params, 1 GPU, and full planning <1 second.
📑: le-wm.github.io
absolutely nuts! crazy speed and latency
A breakthrough in real-time video generation.
As a research preview developed with @NVIDIA and shared at @NVIDIAGTC this week, we trained a new real-time video model running on Vera Rubin. HD videos generate instantly, with time-to-first-frame under 100ms. Unlocking an entirely new creative paradigm and bolstering the foundations of our General World Model, GWM-1.
Real-time generation opens a fundamentally different design space for video models and world simulation. We're investing in co-designing our models alongside advances in hardware to keep pushing this frontier.