@rom1504i
iAccount based inFrance
About this account
- Account based in
- France
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
Gemini Video at Google DeepMind. Better dataset, better infra, and automate it all. Building Physical AGI
Palo Alto, CA
Joined July 2008
- Tweets1.2K
- Following1.2K
- Followers2.3K
- Likes17K
Somehow it's non trivial for LLMs to produce a json representation of this kapla structure and make it back into a visual with three.js. gpt 5.6 sol and Gemini 3.7 flash can't figure it out, they make very approximate reconstructions.
🤖 Made with AI
The amount of autonomy between Opus 5 and Fable 5 is completely different. Opus needs constant pushing to make it keep working, it admits errors instead of fixing them. Fable simply keeps going fix whatever needs fixing.
I would say the level of autonomy is 5x or so.
Yet benchmarks show the same performance? Are there any benchmarks showing this massive gap?
Romain Beaumont retweeted
The most important word here is *ecosystem*. It's not just about having an open-weight model. Open-weight models are a means to an end. To have a truly strong, open ecosystem, we need four critical frontier-level ingredients: open-weight models, open training datasets, open software stacks, and open process knowledge. Few people realize that NVIDIA actually has been pushing beyond open weights by releasing code and datasets for their Nemotron models, which is something open-weight model developers don't do. Marin further opens up the process knowledge - not just how to train one model, but how to iteratively improve and shape a model given particular goals, custom data, and hardware, e.g., how to design scaling laws and evals to guide architecture and data ablations. Open weights, datasets, software, process knowledge: these are the four critical ingredients (renewable resources) that give everyone the ability to most efficiently turn their compute (consumable resources) into the best models according to their needs and values.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
images.nvidia.com/pdf/Open-W…
Supporting open weight is a great first step. When does the research, process and data to make the next models become open again?
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
images.nvidia.com/pdf/Open-W…
I agree with this (and above ideas in the thread). For new code bases we should not need to look at the code, both as an author and a reviewer we should need only to understand the concepts (abstractions, long term plans).
Replying to @matteocollina
Yes but not because I believe it's really needed. I do it because of what the user base average belief is right now. It is at this point a matter of respect and the fact that people will want to modify Redis manually. I can't code Redis thinking just at myself.
I was curious how good gpt 5.6 sol high was so I decided to try out giving it and opus 4.8 xhigh (tried to use fable but it fell back immediately!) a reverse engineer of a small ear camera I had. It works by having an android app connect to the camera over wifi. So I gave the apk to both agents and let them work. 1h later codex had a CLI that could stream the camera to mpv. Claude had picked the wrong path and could not connect to the camera.
So I got a crazy idea today: can codex create a simple Android app based on this reverse engineering of the protocol? I asked it to do that, 30min later it installed it on my phone with adb and it works github.com/rom1504/ear_camer…
This is very impressive.
It means you can basically clone Android apps with a clean room process.
In the process both agents looked at the decompiled java code then at the decompiled assembly from the native drivers and figured out the native protocol from there.
I did that manually a bunch of times over the years for minecraft java at github.com/PrismarineJS/node… and minecraft bedrock (c++) and it's very much non trivial and usually requires lot of try and error so it's pretty cool this is automatable in 1h now.
I think it's quite interesting there is a similar relation between the mass of propellant you need to take and the mass you actually want to move to escape earth velocity and to reach 0.25c.
Said differently, escaping earth and reaching relativist speed are both fairly impossible at large mass.
That's assuming you are using propulsion as your mean of gaining speed... There are some other ways...
🤖 Made with AI
Could be possible to build an automated robotic factory around mercury to build increasingly many solar satellites that put their solar production into anti matter. After 20 years of exponential growth we would have 1000 tons of anti matter which is enough to make a 1000 tons ship (enough for 10 people) go to alpha centauri in 20 years
Opensource degrees for models
1. Open development, everyone can see and contribute to the process
2. Everything is open source except for the process which is closed
3. Weight and code open but the data is closed
4. Only API
5. Only internal usage
6. Only the model can use itself
GLM 5.2 is 3
Opus Gemini GPT are 4
Mythos is 5
Hopefully we don't get to 6!
Would be amazing to have more projects in degree 1, an example outside of models would be linux where not only you get the code and can contribute but you can read the emails discussing the development process and understand *how* linux is being built.
Opensource has benefits and drawbacks but I think it helps to think how open a project is and whether that is the optimal choice.
Humans are still much better at roadmaps and plans than agents.
Why is that?
How are you scaling the planning phase now it's so easy to spawn a project every hour with coding agents ? Agent can't keep track of plans and don't have good taste.
Should people spend most of their time doing planning?
Agents should be able to understand reality.
I have this model of planets with gears. It's a toy for kids so most likely it's not completely correct.
Trying to ask different llms and agents if they can measure the speeds.
Results in next messages.