@ottogin1

helping companies train models @ https://nitter.cf/t.co/MDNJs6T5eA | PhD at MIT CSAIL | SF and Boston

San Francisco
Joined December 2018
Made a skill for Claude to run iterative optimization loops based on using it for 6 month to trainin models for me
1
1
16
2,001
Artem Lukoianov retweeted
Today we’re releasing Adam and Eve, the most human-like AI voices ever built. Adam ranks #1 among AI models on @DesignArena’s AudioRealismBench. We believe we’ve crossed the uncanny valley. API access on our Website! Listen: freyavoice.ai Leaderboard: designarena.ai/leaderboard/a… Special thanks to @alpsencerozturk, @ahmeterdempmk, and the entire Freya team for making this possible. We’re just getting started. Join us!
215
122
45
1,237
485,098
2 days ago Navier Stokes, now this. If we don’t get teleportation by Saturday, I am disappointed
fly brain weights 30ug draws 200nW power, but with proper training, it can drive car and play video games NVIDIA B200 has 17x denser transistors than fly brain(if they are digital), but is 2615x less energy efficient human designed computation systems have a long way to go, big potential!
2
2
4
1,396
TL;DR for me: •RSI take off is very possible and now is main focus of OpenAI •Alignment is getting harder and harder as chain of thoughts monitoring becomes less reliable
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
1
2
276
😵‍💫 can’t wait for P=NP
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
4
135
It looks like Astra is big on 3D and visual understanding. Our benchmarks on iterative optimization are coming soon
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
1
404
Fable 5.1 is surprisingly indistinguishable from Opus 5 on our benchmarks. What is your experience with Fable 5.1?
1
3
375
We already hosted 2 events gathering top people thinking about these questions day to day at companies like OpenAI, Google DeepMind, Meta and others.
1
54
Some great reads to start thinking about RSI: 1. Famous article from April 2025 predicting how AI will evolve in the next few years : ai-2027.com/ and its newer version with the proposal for a better aligment ai-2040.com/ 2. Great literature review from Lilian Weng: lnkd.in/e7tGKqbB
59
We are building a new community around RSI in SF. Come join! We are meeting every single Sunday for a full day to discuss the latest research in RSI, exchange notes and try quick one-day long experiments.
4
14
529
There are still a lot of questions remain that re insanely hard: * How to automatically create evals and RL environments * How to create good rubrics for LLM-as-a-judge and overcome its biases * How to make sure self-improvement is not overfitting and not mode collapsing * How to make the RSI system aligned in case of the take-off
1
1
95
RSI is already happening in some sense as the frontier labs' code is 99% AI generated. If achieved, RSI can be the most important technology and completely change the course of AI development.
1
1
121
The quality is insane!
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
12
930
Real autoresearch is coming.. AI is capable of achieving so much when put into the right feedback loops. And now these feedback loops will be grounded in the real world.
Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. Read more: anthropic.com/news/model-har…
3
356
This is such a cool idea to gamify RL. Expecting to see lots of insane attachments and add-ons that people train it on
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: pollen-robotics.com/microduc… (video with sound on 🔊)
4
339
Yesterday was a great day
yesterday we hosted an incredible hack for RSI & autoresearch in SF with Autolab (@ottogin1 and the group are absolute pros at organising hackathons) we had almost 400 registrations and sadly could not take in everyone, but stay tuned for future hacks. teaser: there might be one just around the corner next Sunday on #RSI & Harnesses
1
8
441
Hot take: this is just a smart way of scaling thinking tokens
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: github.com/llm-as-a-verifier… More on verification scaling in my previous post.
342