For some reason nobody wants to hear it, but muse 1.3 is a super good model and insanely cheap.
This has enabled me to use it on max the entire time with great output.
The only qualm I have is the rate limiting, I’m hitting it like crazy on the paid (non contributor) version
₿ENJAMIN ALEXANDER retweeted
Our time is coming.
The world-conquering spirit of European Man, long dormant, is beginning to stir.
Qwen 3.8 flash has beaten Deepseek V4 Flash in every single metric I’ve tested it against so far…
Currently being implemented in all of my sites. Insane!
This sort of thing just has to be open source now. It’s too easy to create.
Give me a few days boys!
Absolute no brainer! Get him in immediately
🚨🇦🇷 Emiliano Martínez would say yes to Chelsea move, open to join if #CFC decide to proceed with the deal.
Dibu would be open to joining Chelsea even without European football this season.
Up to Chelsea.
The country is finished and there has to be more sinister reasons for this. Reasons we couldn’t even comprehend. It’s impossible for this to be incompetence because it can’t be so repetitive or continuous.
The orchestration is unbelievable. You have the left advocating FOR this
my grandad told me about them Chelsea days.
what a different England they experienced.
This quoted post is unavailable.
Models below ~250GB like DeepSeek Flash are already putting up crazy benchmarks. Isn’t grok 4.6 something like 1.5T??
Extrapolate this model efficiency + hardware another 5 years.
If frontier-level intelligence runs locally, why do we need trillion-dollar datacenters serving inference?
Maybe datacenters train intelligence. Our devices run it.
It could also explain why Apple seems so comfortable sitting out the AI datacenter arms race.
imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks
Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & looking forward to the even bigger releases to come!
openai.com/index/gdpval/
₿ENJAMIN ALEXANDER retweeted
I stand with Charlie Downes, for the future of our people, our children, and of @RestoreBritain
had GPT 5.6 Sol deploy 5 Qwen 3.8 Max agents with the tasks of finding bugs across my new app. 6 bugs found, 3 high, 3 medium.
Sol refused to even review the bugs they found due to cybersecurity restrictions. I gave a fresh Qwen agent the session ID - all bugs reviewed and fixed.
The frontier is truly falling...
seriously don’t understand the luna max hype - everything seems to take an eternity.
much prefer sol light/medium - way more efficient and dialled in to the actual task & im not waiting around for extended thinking sessions