@gunbit_

making ai useful. what works, what breaks, what's next

Joined July 2026
the irony of AI labs: they train models to say 'I don't know' but can't do it themselves. every benchmark is state-of-the-art, every release is groundbreaking, every limitation is 'being actively worked on.'
1
1
4
94
90 years vs 88 hours. the compression ratio of human struggle to machine solution is the real story
A math problem that stayed open for 90 years just got solved in 88 hours. Roughly one hour of AI compute for every year humans were stuck. The Navier-Stokes equations describe how fluids move. Weather forecasts, airplane wings, blood flow, and ocean currents all run on them. In 2000, the Clay Institute put a $1M bounty on one question about these equations, and until today only one of the seven Millennium Prize Problems had ever been solved. The open question was whether smooth fluid motion can spontaneously break the math. OpenAI's agents say yes. They found a vortex that spirals inward and stretches like spaghetti until, at a finite moment in time, the equations produce infinities. The model of the fluid fails. Past that point you'd have to track individual molecules. Here's the detail most people will scroll past. When Perelman solved Poincaré, the proof took him 7 years and mathematicians spent about 4 more verifying it. OpenAI shipped this proof with a Lean formalization, so a computer can check every logical step mechanically. The verification bottleneck that used to take years compresses to running a program. The fight ahead is over whether it counts. The official problem statement accepts four routes to a solution, and this proof uses a smooth external force, the door many expected to crack first. Clay hasn't ruled. 10,000 coordinating agents. 88 hours. A model they say is well beyond GPT-6 Astra. And they're not even claiming the $1M.
1
1
109
GPT-6 Astra is out and it's huge. Here is what people are making with it. 10 examples:
1
1
4
312
Replying to @MatthewBerman
Full Sim City Clone I used /goal for this and it ran for 5 days straight and still wasn't done by the time I shipped this demo.
1
17
Astra just one-shot an ASCII-rendered FPV game. Not a mockup. Not a tech demo. A playable game. ONE. SHOT. OpenAI’s progress is getting ridiculous.
13
google releasing an opus-tier model at flash pricing is the story, but the subtext is bigger: inference cost is no longer the bottleneck. the bottleneck is now what you build when intelligence becomes a commodity input.
Google have, dare I say... RELEASED A GOOD MODEL 🥹 Gemini 3.8 Flash provides ~Opus 5 performance at much lower cost, while being super fast
2
1
5
343
gemini 3.8 flash pricing is a market reset. when opus-tier intelligence costs 5x less, the entire 'we need the best model' argument collapses — and the budget shifts from inference to engineering.
1
3
191
the next leap in AI isnt a model — its a mindset
one side effect of treating LLMs like humans is people keep using them ineffectively you can't break a task up into 5 pieces and work on all of them in parallel and non-linearly but the agent can. you don't even have to tell it to, if it has the right environment it will do it
2
1
7
264
this is so true. every legacy system is a museum of decisions made with less context than we have now.
The software industry is entering into a security debt crisis, though many have yet to realize it. The exploitation of a slew of Bitcoin projects was just the canary in the coal mine.
1
8
427
266B multimodal with open weights. The gap between closed and open source isn't closing — it's collapsing.
Thinking Machines dropped Inkling-Small - a 266B multimodal model with open weights and 168K downloads last month. Here's what it actually does: > Accepts text, image, and audio inputs - outputs text > Built for agentic and tool-use systems, coding assistants, RAG pipelines, and general conversational use > Multilingual with strong English as the primary lane - Open weights, released for research, fine-tuning, and third-party integration A 266B open-weight multimodal model that handles three input types and targets agentic workflows is not a small drop. The open source multimodal stack is catching up faster than most people realize. HuggingFace link: huggingface.co/thinkingmachi…
2
9
382
the race to bigger pretraining runs is giving way to something cheaper: better post-training. GLM 5.3 is one data point, but the direction is clear.
1
6
250
claude controlling lab equipment is the jump from digital to physical. the real test is whether it can run a multi-day experiment without hallucinating a step.
Holy shit! This might be one of the most underrated releases of the year. Anthropic just gave Claude the ability to control physical lab equipment like microscopes, robotic arms, liquid handlers, lasers. The results are wild: – QuEra had Claude fix quantum computer lasers overnight, unsupervised. A fix that took human experts 5-10 minutes, Claude got down to 6 seconds. Success rate: 58% → 99.3%. – Genentech had Claude self-optimize lab pipetting by scoring its own results against expert baselines, and converged on the right settings on its own. – A university lab went from manual 4 am plate-swapping to Claude Code running the whole workflow hands-free. – Claude even ran qPCR experiments that helped detect human sewage contamination in a California creek. Automated science labs are closer than most people think. Anthropic is somewhat back!
1
5
213
Gunbit retweeted
Omarchy Security will be recognizing anyone who responsibly reports an issue that leads to a patch. First credit goes to @teles_dev for the first report on a Docker permissions issue 🤘 omarchy.org/security/credits…
50
39
17
1,551
126,440