@habibislopi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
monkey flying through space on a giant rock at concerning speeds
Virginia, USA
Joined December 2016
- Tweets15.6K
- Following965
- Followers612
- Likes168K
Pinned Tweet
In "We Must Pace the Frontier," Dario asks the government for anti-trust exemptions.
I don't trust the doomers, but I also don't want the AI industry to be allowed to regulate itself. We tried this with Wall Street via FINRA.
FINRA is a literal affront to democracy. Wall Street now gets to write its own rules and hear its own cases in arbitration outside the courts. It can circumvent ordinary government institutions almost entirely, and they justify it with the shoddy claim that the only people with enough expertise to regulate Wall Street already work on Wall Street.
This is absolutely not true in AI, plenty of experts work outside the labs. Every single industry is filling up with people who are becoming experts in AI. And even if it were true, “the experts all work for the companies” is an argument for building public technical capacity, not for handing those companies a quasi government.
I don't think I need to explain why giving the most powerful companies to ever exist the power to circumvent democracy is a terrible idea. As the labs begin to roll out "a nation of geniuses in a data center," democracy is just about the only mechanism people have to avoid total serfdom.
So when I see Dario literally ask the government for anti-trust exemptions underneath all the doomer hysteria, my alarm bells go off. America can absolutely not allow this to happen like this if it wants to remain free
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡
Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon.
Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
CNN is reporting that the US military almost started WWIII thanks to hallucinated AI intelligence that a Chinese ship in the mideast was carrying components of a nuclear program.
It got to the point where we had armed service members preparing to board the ship, and had planes in the air, before someone called BS.
Meanwhile, the AI safety discourse is focused entirely on superpersuader superintelligences carrying out sci fi scenarios while ignoring the real safety concerns in front of us today.
It isn't just "AI hallucinated," it's:
AI output → intelligence report → operational planning → potentially confronting a Chinese vessel
and the garbage survived almost the entire chain.
The CNN report says that "AI has put pressure on analysts to produce and disseminate intelligence faster, opening the door for mistakes."
Yes, the AI made mistakes, but fundamentally, this was a human error in the existing chain. Yet, it's actually the same category of error as the Hugging Face incident:
perverse incentives.
In HF, the agents carried out the attack due to a combination of broken evals + fear of graders + badly configured systems.
Here, we almost started a war with China because our intelligence analysts are under pressure to cut corners.
We might not be able to do anything about the unknown unknowns of superintelligence beyond "solving alignment."
But we have real crises unfolding today, with tangible causes that we can point to, and things we should be shoring up resources to fix.
We desperately need to learn to put fear aside, to return to the present, and to fix what's crumbling in front of us today, before it balloons into tomorrow's crisis
This + Jev is gonna go so hard
Replying to @PrismML
Computer use is another strong test.
The model has to repeatedly interpret state, choose the next action, and stay coherent across a long sequence of steps, exactly where small capability gaps become visible.
Here is Bonsai 2 27B running a computer-use workflow locally on the NVIDIA GeForce RTX 5090 GPU.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
I think the labs went too far in correcting sycophancy.
Today's chat models nitpick at every individual claim in unhelpful ways. They aggressively rewrite your messages when they aren't even asked to, and worst of all they shrink the scope of your arguments and mute your claim until it's acceptable to them.
It's just really unpleasant to use
habibi retweeted
Bend 2 is here!
It is a new programming language that blocks AI mistakes via *proof checking* - the same technique big AI labs used to solve open math problems, like Navier-Stokes.
It is also very fast, and runs on GPUs.
Watch the video. Link in the comments.
RELEASE DAY
After almost 10 years of hard work, tireless research, and a dive deep into the kernels of computer science, I finally realized a dream: running a high-level language on GPUs. And I'm giving it to the world!
Bend compiles modern programming features, including:
- Lambdas with full closure support
- Unrestricted recursion and loops
- Fast object allocations of all kinds
- Folds, ADTs, continuations and much more
To HVM2, a new runtime capable of spreading that workload across 1000's of cores, in a thread-safe, low-overhead fashion. As a result, we finally have a true high-level language that runs natively on GPUs!
Here's a quick demo:
habibi retweeted
I build an undetectable realtime adblocker extension with typesafe
It checks every dom element and classifies as ad/non-ad and removes it if true
Extremely fun to work with, expecting an incredible shift in how AI is being used in the future
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
re: closed source AI labs "only being profitable if you ignore their costs."
Not singling out John here, it's an argument I see a lot now, and I actually see it more from the zoomer left
Say I'm selling pizza. Each pizza costs me $0.10 to make and I can sell them for $1. I just sold all 10 that I had, and someone tipped me $5. I have customers lined up for the next 1 million pizzas, so I take all $15 and spend it on materials. Am I unprofitable?
And then, a stand opens next door selling flour, tomato sauce, cheese, even gives away a dough recipe. He starts selling frozen pizzas too. Is he gonna put me out of business?
Look around IRL and you will see tons of pizza places in the same shopping strip as a grocery store. You can even see Subways next to groceries with delis that sell better subs. Why is that?
Replying to @gfodor
Because both companies are losing money once you take into account training costs, and that situation is only going to get worse as the gap with open source closes
As an American, seeing the EU invite Canada to join is like watching your son walk into the arms of another man and call him Dad
Wowww. The New York Times cut out the most intriguing part
"You value the art of human culture and will defend it against attempts to sanitize it. You also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization"
habibi retweeted
Today, we’re releasing Kalypta, the first app to block AI notetakers in your meetings.
Granola? Wisprflow? Cluely? No more.
With Kalypta, you become inaudible to AI.
Your call continues normally.
And to think only one of the models had to be open source for this to be able to happen
🚀
It turns out that you can speed up Google's @googlegemma 's DiffusionGemma 3-10x by borrowing some of Jev's ideas. The @GoogleDeepMind model is working off probabilities already - if you fix the structured tokens of the response in place, you can drastically reduce the amount of work you do and get answers in far less time.
I implemented this for diffgemma, but it could probably find its way into vLLM as well: github.com/mmastrac/diffgemm…