@armeencoi
iAccount based inCanada
About this account
- Account based in
- Canada
- Connected via
- Canada App Store
Account-level information from X, not a live location or the device used for a specific post.
Co-founder, Comox AI. Eval data for VLMs / multimodal models. armeen@comoxai.com
Vancouver
Joined May 2026
- Tweets267
- Following13
- Followers4
- Likes225
At Comox we build eval data for VLMs and multimodal models. Frontier models score under 20% on our hardest tests, and 0% when they have to count objects exactly.
Curious if that kind of data is useful to your team. I can send the breakdown.
founder@comoxai.com
🤖 Made with AI
Hard eval data for VLMs and multimodal models
Frontier VLMs are often good enough on clean benches and still fail on messy real footage, exact object counts, rare conditions, long-tail scenes public data and sim cover poorly. Useful ≠ accurate under those conditions.
The line that matters here isn't "world model," it's pixel-perfect camera control. Generation you can't steer is a mood board. Generation where you pick the camera is a simulator, and that's what robotics has been waiting on.
🤖 Made with AI
The headline here isn't the fidelity, it's "pixel-perfect camera control." That's the difference between a world model you watch and a world model you query. Same scene, new viewpoint, on demand, is how you get eval sets and training data that aren't just your own lab from one angle.
🤖 Made with AI
World Labs shipped Atlas today: a multimodal world model with explicit camera control that reconstructs generated frames in 3D. The interesting claim isn't the visuals, it's controllability — geometric consistency under a specified camera path is what makes generated worlds useful as training and evaluation environments rather than demos.
🤖 Made with AI
Vendor hours were never going to get you there. Owning capture is the move. The remaining gap is held-out eval, not more video.
🤖 Made with AI
Hydra-0 shipping as an open model with action flow in pixel space is the real news here: a shared action interface other labs can actually test against. What's still missing is a held-out scene bank to score it on. nitter.cf/NVIDIARobotics/status/…
🤖 Made with AI
Robot actions can be represented as motion in pixel space.
NVIDIA Research introduces Hydra-0, a generalist world model conditioned on action flow: image-plane trajectories that enable one model to learn across human hands, handheld grippers, single-arm robots and bimanual systems.
Explore the project 📄 nvda.ws/4xaO9C0