I mostly tweet about #ai, #robots, #science, @packers... Senior SWE at @Anduril | Ph.D. Robotics @GeorgiaTech | M.S Robotics @Penn Thought/opinions are mine

Atlanta, GA
Joined October 2012
Jonathan Balloch retweeted
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia? These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach. The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date. The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant. At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems. So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models. Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field. If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section. Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models. I'm interested in your thoughts.
91
390
67
2,368
677,311
Jonathan Balloch retweeted
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing. - pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper - use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this - use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs" "De-identification" is weak -- you can identify someone with a small number of bits, and long traces have more than enough. And it doesn't affect IP leakage concerns.
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
44
154
34
1,361
282,876
This thing is so so cool
You can now play this for real with anyhumanever.com/. I just died of smallpox aged 7 in 1889.
2
41
Jonathan Balloch retweeted
You can now play this for real with anyhumanever.com/. I just died of smallpox aged 7 in 1889.
People generally have no idea how much the past sucked. It's fun to study history but anyone who genuinely romanticizes the past I struggle to take seriously. Even the recent past was miserable by comparison with today, but the pre-industrial era was a total horror show on every axis continually. Reply here with your hypothetical pre-modern name, location, and year of birth and I'll give a quick summary of how your life went. I'll start with a historical character, my 5th great grandfather John Jones, born ~1812 in northern Wales. Illiterate. Orphaned. Laborer. Married at age 33 (already ahead of 75% of other men in that era), 12 children, all but 2 died in childhood, moved (voluntarily) to Australia in 1849, financially ruined at least twice, alcoholic, wife died suddenly aged ~45, unable to care for surviving children, who variously ran away, ended up in some hellish convict era Australian orphanage, drowned, etc. Disappeared in custody c. 1860.
71
86
34
848
155,375
Truly the most important lesson in robotics is that people want solutions to their problems, not robots
Replying to @mehul
1. Customers want solutions to their problems, not robots. Robots are just a means to an end of solving those problems. 2. 'Make something people want' is great. But, in hardware, you must make something people need, not just what they want. Because you have one shot.
1
71
Top 3 launch video. Easy
Agents f*ck up. We raised a $2.3M pre-seed to warn you before it’s too late
53
Jonathan Balloch retweeted
“we sandboxed the agent” meanwhile the agent:
261
3,431
247
35,642
1,749,980
Yup
"let me [...] rather than guessing." There is not a single turn with AI coding without this. Is it because labs haven't found a better solution to hallucinations than littering their system prompt with various paraphrases of `READ THE ACTUAL DATA YOU MF DON'T MAKE IT UP`
111
Perhaps the highest value free resource on the internet
My summer project is done! A 20 video, free course on post-training to accompany my book is all on YouTube with slides open for modification & re-use. ~12 hours of content covers the core foundations and some research areas I think will grow in importance. It was a fun time to review all the fundamentals again, as it is clear in the next 1-3 people the amount of people wanting to learn post training will likely 100X again from today, as we have already 100X'ed from two years ago. As AI agents get increasingly capable at coding and discussing these fundamentals (see the code exercises accompanying the book that I am refining with the community) I think developing clear intuitions for how models work and why is one of the most important skills going forward in AI. Still, learning the post-training math is the best way to battle test them. I personally just in this course am starting to master how forward/reverse KL relates to post-training topics. Thanks to all my viewers, and I'm happy to answer questions in the book discord or understand how to better teach the various reward models, on-policy distillation, new RL algorithms, etc. Plus, the book is 50% off right now with the code PBLambert on Manning to celebrate the launch. I'll share the relevant links below. Who's going to make this course for pretraining?
3
274
I guess some people need to learn they are wrong the hard way
I’m pretty sure gaming as a category is dead because everyone is just gonna build their own games Games don’t have ongoing maintenance cost like companies replacing SaaS vendors Perhaps an argument for wanting someone else to build an experience for you, easier to watch a movie than create one. But video games and vibe coding are so close the difference between creating the form and playing in a form is getting closer and closer to zero. So why play someone else’s game? Just pay chatgpt $20 and make you own
104
Nice! thats doable!
Replying to @mcuban
For 2k-5k sq ft at $25-55/sqft ops (excl personnel): $50k-$275k fixed costs/year. At 15% margin, breakeven sales = fixed / 0.15 → $333k to $1.83M annually. (NYC rents push the high end up further. Labor would raise it more.)
58
Jonathan Balloch retweeted
Never mind the ethics of AI use, we need to establish the etiquette. Inviting me to read LLM-generated text without telling me that's what it is, as though you'd written it yourself, is _rude_
31
289
12
3,394
49,217
Crazy how people will be "disrupters," but then god forbid a *politician* try to do things differently instead of a Stanford dropout tech bro and they suddenly are so pro-establishment that they and their comment section bubble not only predict failure but actively hope for it
These folks have no idea how to run a business, have never run a business, and — wait for it — they put out an RFP to let entrepreneurs run their state sponsored grocery stores So… the government is going to pay companies to discount groceries, which have a 1-2% profit margin Amazing!!! 🍿 Can’t wait to see them use some insane DEI criteria to select the RFP winners and this entire experiment to devolve into chaos! COSTCO ALREADY HAS ZERO PROFIT MARGIN GROCERIES — YOU MIGHT AS WELL JUST NEGOTIATE BUYING ACCESS TO THEIR SELECTION FOR DISADVANTAGED NEW YORKERS!!!
46
Cortana experiencing rampancy @jamessealesmith @andrewsilva9 @cusuh_
when I have to kill the long running session that has all the good info and a constructive back and forth because claude just starts getting context window dementia
3
120
Atypically bad take from karpathy. Its not hard for LLMs to generate *something*. if you expectations are limited to (a) existence and (b) polish, sure, we have reached "excellence". but the other day I needed a to vectorize a freehand drawing and VLMs still couldn't hack it
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all. I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand. Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
90
Jonathan Balloch retweeted
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-W…
16,037
29,401
10,948
171,957
66,424,568
Jonathan Balloch retweeted
at this point the distillation arguments need to die and understand that China is also very good at building models
32
46
6
1,795
76,160
Jonathan Balloch retweeted
It's obvious that eventually a speedrun for RL will stick. I currently think the biggest bottleneck is price, as a individual entry currently has too much noise from instability of RL, so running multiple seeds makes it cost O($100). Glad to see attempts!
With RSI around the corner, it's time for an RL speedrun. Introducing Sokoban Speedrun: training Qwen3-4B-Instruct with RL to solve Sokoban puzzles. We start by modding @karpathy’s nanochat RL pipeline; the GRPO baseline takes 87 minutes on 8×H100s. 1/
4
10
1
164
33,654
Wtf this is crazy lol
Replying to @juliarturc
The fact that he praises a model that literally hinders AI research is indeed the end of an era...
60