@pmigdali
iAccount based inPoland
About this account
- Account based in
- Poland
- Connected via
- Poland App Store
Account-level information from X, not a live location or the device used for a specific post.
Crafting challenges for AI - founding engineer @QuesmaOrg. AI & data viz specialist with PhD in quantum physics. Blogs about tech, neurodiversity and stuff.
Warsaw, Poland
Joined February 2011
- Tweets3K
- Following992
- Followers2.3K
- Likes4.3K
Navier-Stokes Equation, explained with colors
p.migdal.pl/equations-explai…
> Make in three js (pnpm) visualization of all Invisible Cities by Italo Calvino. Don't ask questions, it is a one-shot task. You have 6h of work, use it until it becomes a masterpiece.
GPT-6 Astra: p.migdal.pl/invisible-cities…
Claude Opus 5.5: p.migdal.pl/invisible-cities…
I asked Opus 5.5 to explain camera focus by building an interactive lens lab
Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost
lens.lab.sael.net
Move the focus ring and you can see the glass elements shift the sharp plane through the scene
MiMo-V2.6: The Hard Road to Scaling Up RL
MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems.
The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in.
Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category.
There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time.
To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward.
MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
GPT-6 Astra rocks at puzzles - 3d game Portal, Baba Is You, ARC-AGI-3, MazeBench.
Is there a puzzle game still to hard to for this model?
Or maybe all that remains is unsolved historical codes?
quesma.com/blog/gpt-6-astra-…
Quoting @FakePsyho notes on PuzzleScript and ARC-AGI-5 and @patience_cave on Fable 5.1 and Astra results MazeBench.
Piotr Migdal retweeted
Using Claude in Chrome product. Prompt:
replicate this image (my profile picture) in jspaint, use brushes to replicate it as closely as possible. I want you to use brushes and clicking. Do not use javascript tool, but the final result has to be very close to this image. DO NOT CHEAT
1. I specified the do not use javascript tool because otherwise it uses javascript and it is faster but it seems like people cared more about clicking :D
2. The task took 45 minutes, but I don't think it's that much slower than competitors 😄
3. Prompt is important. (See thread)
Piotr Migdal retweeted
GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5.
OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.
Piotr Migdal retweeted
I threw a ton of various logic puzzles at Astra and... it logically (no code) solved ALL of them. I mean literally every puzzle that I tried, including: too hard for world championship, very large grids, unpublished puzzles, puzzles that are impossible to solve via backtracking alone.
For comparison, sol 5.6 was around 20-30%.
I analyzed the results and all of these were legit.
Honestly I don't care if OpenAI RLed logic puzzles to death. Astra is now better at explaining solving paths than I am.
Piotr Migdal retweeted
And... GPT-6 Astra has autonomously completed Portal! I didn’t expect this to happen so soon, but I’m glad we've made so much progress here.
I was reminded that back in 2016, one of OpenAI’s technical goals was to “solve a wide variety of games using a single agent.”
Piotr Migdal retweeted
since you guys loved the exploding tesla..
I used GPT-6 Astra to create a 3D website that pulls apart the male anatomy into 2,234 modeled pieces!
we are in a renaissance of learning
Only GPT-5-Astra and Claude Fable 5.1 solved all levels from "The Lake" stage of Baba Is You.
Astra is 2x faster and 3x cheaper.
Bench (now saturated): quesma.com/benchmarks/babais…
If you go mushroom hunting, do not rely on AI - it can be deadly!
Even frontier models, while good at identifying many mushroom species, still have a dangerous error rate, calling poisonous mushrooms edible.
Blog post and benchmark: quesma.com/blog/mushroom-llm…
Which local models are the best at drawing pelicans?
Even if we go to the extreme, and have 128GB of RAM.
Qwen3.8 Flash-Next gives a lot of details.
DeepSeek V4 Flash is strangely underwhelming.
Qwen3.8 27B still rocks, and I like its consistent minimalism.
Is Qwen3.8 27B still large at 31GB?
It is! But for this tasks 2-bit quantizations (at around 12GB) will give the same results. For more complicated coding, 4-bit are more than enough.
quesma.com/blog/qwen38-27b-q…
4 bit is all you need - at least when it comes to Qwen3.8 27B, running an agentic coding benchmark Terminal-Bench 2.1.
quesma.com/blog/qwen38-27b-q…