PhD student in robotics manipulation, currently on physics-aware world models for robust manipulation @RiceCompSci. Designer and developer of XLeRobot.
Houston, TX
Joined November 2021
- Tweets505
- Following305
- Followers2.2K
- Likes1.3K
Pinned Tweet
Catching a flying ball is hard. What if with a flat plate? 🏓
Our work at RSS’26 shows it’s possible through Zero-Shot Sim2Real
With Domain-Randomized Instance Set (DRIS), we catches different kinds of balls without any real-world fine-tuning
🔗 rice-robotpi-lab.github.io/D…
Vector Wang retweeted
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids.
- learned directly from human manipulation data
- no teleop/robot data
- close to human-level dexterity and efficiency
- multi-robot collab
Presenting RoboTok: making human skills on the web searchable for robots. Given one demo, it finds similar actions using body-relative 3D hand trajectories across viewpoints and scenes for better dexterous policies.
Project: rice-robotpi-lab.github.io/R…
Code: github.com/Rice-RobotPI-Lab/…
Data and models: huggingface.co/Rice-RobotPI-…
Thanks @NVIDIAAI for supporting this work through the #NVIDIAGrant!
Vector Wang retweeted
🤖 Introducing RoboTok, the “TikTok for robots.”
Just as TikTok recommends videos to people, RoboTok recommends relevant human demonstrations for robot learning.
Robot learning needs broad and diverse demonstrations, but collecting robot data is expensive. RoboTok is an internet-scale data engine that uses web video as a scalable and continuously growing source of demonstrations for dexterous manipulation learning.
Given one human demonstration video as a query, RoboTok retrieves other web videos with similar underlying manipulation motions.
Rather than matching videos by labels or visual appearance, RoboTok compares how the hands move over time.
💡 The key idea is to represent each video with canonicalized 3D hand trajectories. Each trajectory is expressed in an estimated actor-centered reference frame, so the movement is described relative to the person rather than the camera. This makes similar manipulation behaviors easier to compare across different viewpoints and scenes, even when the actor is partly occluded.
In our experiments, RoboTok retrieved more manipulation-relevant demonstrations than existing robot-data retrieval methods. When those videos were used to guide robot training, the simulated robots completed manipulation tasks more successfully.
I’m sincerely grateful to Howard Qian and Kaiyu Hang for leading this project. I also want to thank Yiting Chen, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen @bowenwen_me, and Chen Wei @_Chen_Wei_ for their guidance and collaboration.
Project site: rice-robotpi-lab.github.io/R…
Paper: arxiv.org/abs/2609.03199
Code: github.com/Rice-RobotPI-Lab/…
Data and models: huggingface.co/Rice-RobotPI-…
In the “TikTok” for manipulation data, the searching algorithm should not only be based on the task text description, but also the hand motion itself
Robot learning has a data problem: the internet contains billions of human manipulation videos, but most robot datasets cover only a tiny fraction of real-world behavior.
RoboTok turns those videos into a searchable source of dexterous manipulation demonstrations. Give it one human demonstration, and it retrieves other videos with similar hand movements.
Come and show your works and demos in agentic robotics! See you soon at Austin!
🚨 Call for live demo / paper at CoRL 2026 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗥𝗼𝗯𝗼𝘁𝗶𝗰𝘀 𝗪𝗼𝗿𝗸𝘀𝗵𝗼𝗽! No paper required for demo track. Due Sep 27.
🔮 The real magical power of agent is live interactive demo with the audience, which is why we provide all-out support:
- Robot, compute, and API are ready for you in the conference room
- Custom robot / sim demo also welcomed
- Just submit a video proof - friendly to industrial participants
agentic-robotics-workshop.gi…
Paper submission also welcomed!
We also have an amazing speaker lineup from CMU, UC Berkeley, DeepMind, NVIDIA, and Tencent.
Vector Wang retweeted
We're hiring! The Seattle Robotics Lab (SRL) at NVIDIA is hiring a research scientist to work on robot foundation models. The focus will be on investigating the science behind building VLAs, WAMs, and world models for robotics, and leveraging the findings to propose the next generation of models and training paradigms.
We're seeking a world-class researcher (PhD grad) or senior researcher (3+ years of post-PhD experience) with expertise in training generalist policies. The researcher should aim to question conventional wisdom, push beyond the status quo, and dive into the foundations of general-purpose robot intelligence.
NVIDIA SRL provides a rare environment to conduct open research on fundamental and applied robotics, with the freedom to publish, collaborate broadly, and tackle big scientific questions at scale.
If this sounds like a strong fit, please reach out to us and apply here!
nvidia.eightfold.ai/careers/…
nvidia.eightfold.ai/careers/…
Congrats Stone! The right person to the right place.
Just over a month later, I have now joined @physical_int full time! Wrote a bit about why I'm excited to research simulation at Pi
stoneztao.substack.com/joini…
A faster way to get actions directly from JEPA world models! @ylecun
5*168min=14hrs
We just hit a weird milestone: our model became more reliable than your average home WiFi.
Just like everybody else, we thought cloud inference was the obvious choice. Yet 2 days into the ACT-2 eval, our mind completely changed.
If our hero @ArpitKalla didn’t cook, this video wouldn’t exist 🧵
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Just got message it will be able to be directly SSH into. Can’t wait to let my agent run this. Ordered
A year ago, when I left Tesla Optimus to co-found Mondo, I made a promise to my son: I would build him a little robot buddy — something like the robot friends from our bedtime stories, but real.
At the time, it was not really a product statement. It was a father’s promise to a little boy I love more than anything — a boy who loved stories about a child and his robot companion. Over the past year, that promise slowly became real — Beni the robot.
Beni follows you, films you in 4K, moves across different terrain, jumps over obstacles, captures moments from new angles, and helps turn your clips into highlights. We designed it for families, pets, creators, athletes, and anyone who wants a robot that feels less like a gadget and more like a little sidekick.
I’m deeply grateful to the Mondo team for turning countless prototypes, failures, late-night tests, and difficult tradeoffs into something people can finally meet. Although we have only just launched on Kickstarter, this world-class team is already deep in the “production hell”, and we expect to ship the product in the next few months.
I’m also especially thankful to my family. Building hardware takes time, patience, and sacrifice from the people closest to you. Beni would not exist without their love and support. My son often visits our office and plays with Beni. Watching them together has reminded me why I started this journey in the first place. I do not want him to grow up thinking of robots only as distant machines in labs, factories, or videos. I want him to grow up in a world where robots can feel close, friendly, safe, and present — something he can play with, trust, and remember as part of his childhood.
For the robotics community: Beni may look cute, but it is a serious robotics product. We used reinforcement learning to train Beni’s motion policies, with NVIDIA Isaac Lab and Google MuJoCo as part of our training and validation pipeline. Under the playful exterior are real robotics problems in locomotion, perception, control, and sim-to-real transfer. The lessons we are learning from Beni also directly inform our future humanoid product. Different form factor, same belief: advanced robotics should not feel distant or abstract. It should become close, present, and part of everyday life.
For my son, Beni began as a promise. For Mondo Robotics, it is just the first step.
Meet Beni: mondorobotics.com/
Vector Wang retweeted
"A parcel with snacks has been delivered for Flexion. Retrieve it using the stairs and come up using the elevator. Then unpack it and place the items into the empty drawer on the shelf in the snack area."
One instruction. No human operator. Everything that follows is autonomous.
Today we're introducing Reflect v1.0, our robotics intelligence platform for long-horizon work. From a single natural-language command, the robot understands the task, navigates a multi-floor building, calls elevators, handles doors, uses tools to unpack a box, and puts the items away. The biggest shift in v1.0 is that we use reinforcement learning across every layer, from low-level control to high-level reasoning.
Long-horizon autonomy is unforgiving. The robot must recover on its own when things don't go to plan because in the real world, they never do. Combining reasoning, perception, physical execution and runtime robustness into a single mission-capable system is the foundation required to solve humanoid autonomy.
Our team is just getting started.
#HumanoidRobots #Flexion
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Real world setup even lol
today in “things I didn’t expect people to run in my sim”: sheep getting herded by XLerobot robots @VectorWang2
Vector Wang retweeted
The most inspiring thing I took from this paper: there's far more to squeeze from simulation than sim-to-real training of task-specific policies. RATs shows a coding agent can self-propose tasks, self-construct scenes in sim, and acquire skills that transfer to real-world deployment.
It's promising to imagine handing coding agents a bunch of simulation clusters on top of ENPIRE to enable Sim-and-Real Co-research, where agents massively learn skills and try ideas in sim while continuously grounding them in the real world. Then robot skill acquisition can really scaling like everything else in the deep learning era.
Children learn from play. Can robots do the same?
We propose 𝐏𝐥𝐚𝐲𝐟𝐮𝐥 𝐀𝐠𝐞𝐧𝐭𝐢𝐜 𝐑𝐨𝐛𝐨𝐭 𝐋𝐞𝐚𝐫𝐧𝐢𝐧𝐠, a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with 𝐑𝐀𝐓𝐬 (Robotics Agent Teams), where robots discover reusable skills through curious play.
Co-led with @jiaxin_ge_
Introduce EgoInfinity:
a web-scale (14.6 yrs, 142M clips)
data engine that automatically lifts Youtube videos into 4D hand-object interaction, and retargets for robot learning
Not only a dataset, but a modular, upgradable engine
HF Space: huggingface.co/spaces/Rice-R…
The dataset is open to anyone to use and contribute.
The pipeline is modular and upgradable.
Come scale things up even without any hardwares!
Project Page: rice-robotpi-lab.github.io/E…
Paper: arxiv.org/abs/2606.17385
Codes: github.com/Rice-RobotPI-Lab/…
Dataset: huggingface.co/datasets/Rice…