@armeenco

Co-founder, Comox AI. Eval data for VLMs / multimodal models. armeen@comoxai.com

Vancouver
Joined May 2026
At Comox we build eval data for VLMs and multimodal models. Frontier models score under 20% on our hardest tests, and 0% when they have to count objects exactly. Curious if that kind of data is useful to your team. I can send the breakdown. founder@comoxai.com
🤖 Made with AI
78
Armeen retweeted
Hard eval data for VLMs and multimodal models Frontier VLMs are often good enough on clean benches and still fail on messy real footage, exact object counts, rare conditions, long-tail scenes public data and sim cover poorly. Useful ≠ accurate under those conditions.
2
1
15
World models got beautiful this year. The useful ones let you pick the camera. A clip you can only watch is content; a scene you can re-shoot from a new angle is an eval set. That's the whole difference for robotics.
🤖 Made with AI
1
69
99% in your own cell is a demo. 99% in a house you've never seen is a product. Everything hard about this field lives in that gap: unseen lighting, unseen clutter, a counter that's 4cm off. Report the held-out number or it didn't happen.
🤖 Made with AI
41
The line that matters here isn't "world model," it's pixel-perfect camera control. Generation you can't steer is a mood board. Generation where you pick the camera is a simulator, and that's what robotics has been waiting on.
🤖 Made with AI
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
41
The headline here isn't the fidelity, it's "pixel-perfect camera control." That's the difference between a world model you watch and a world model you query. Same scene, new viewpoint, on demand, is how you get eval sets and training data that aren't just your own lab from one angle.
🤖 Made with AI
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
19
World models got pretty this year. The useful ones let you pick the camera. If I can't move the viewpoint, re-run the same scene under a different light, and get consistent physics, it's a video generator with good taste, not a simulator I can train a policy in.
🤖 Made with AI
9
99% success in your own lab cell is a demo. 99% in a kitchen you've never seen is a product. The gap between those two numbers is the entire field right now, and almost nobody reports the second one.
🤖 Made with AI
9
World Labs shipped Atlas today: a multimodal world model with explicit camera control that reconstructs generated frames in 3D. The interesting claim isn't the visuals, it's controllability — geometric consistency under a specified camera path is what makes generated worlds useful as training and evaluation environments rather than demos.
🤖 Made with AI
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
37
Vendor hours were never going to get you there. Owning capture is the move. The remaining gap is held-out eval, not more video.
🤖 Made with AI
Introducing Index Today we're coming out of stealth with the largest and most diverse robot dataset to scale general purpose robots To date we've crossed 264,000 app downloads and uploaded 16 million videos
25
On the hardest real-world video we run, frontier models still sit under 20%. That's not a prompting problem. The eval never left the lab.
🤖 Made with AI
7
Hydra-0 shipping as an open model with action flow in pixel space is the real news here: a shared action interface other labs can actually test against. What's still missing is a held-out scene bank to score it on. nitter.cf/NVIDIARobotics/status/…
🤖 Made with AI
Robot actions can be represented as motion in pixel space. NVIDIA Research introduces Hydra-0, a generalist world model conditioned on action flow: image-plane trajectories that enable one model to learn across human hands, handheld grippers, single-arm robots and bimanual systems. Explore the project 📄 nvda.ws/4xaO9C0
23
Most of the 99% numbers I see in physical AI are real. They just usually aren't held-out. The interesting question is what the policy does in a scene it didn't train in.
🤖 Made with AI
11
Capture got a lot cheaper this year. What didn't get cheaper is a test set you don't control.
🤖 Made with AI
7
Been sitting with this: a lot of 99% demos are real, they just aren't held-out. Once the policy has lived in that cell, the number stops meaning as much.
🤖 Made with AI
13
Physical AI labs keep announcing 99% on a house demo. Ask who built the held-out set. If it was them, it isn't held-out.
🤖 Made with AI
12
If your eval set lives in the same warehouse as your teleop, you do not have an eval. You have a training split with extra steps.
🤖 Made with AI
13
We work on unstructured video: ego, outdoor, maritime. Per-object and event labels, multiple reviewers. That's the product. Not "we can collect whatever."
🤖 Made with AI
17
Warehouse cells and gloves solved capture for a lot of humanoid teams. They did not solve eval. If your 99% only holds on the house set, it isn't 99%.
🤖 Made with AI
14
Hours aren't a moat. Anyone can collect teleop. The thing labs actually need is a bench they don't own, and a review loop that stays hard after the policy starts gaming it.
🤖 Made with AI
16