@pmbstuffi
iAccount based inCanada!
About this account
- Account based in
- Canada
- Connected via
- Canada Android App
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Full Stack AI Engineer & Researcher. Building AI-powered tools.
Canada
Joined July 2010
- Tweets1.5K
- Following208
- Followers182
- Likes77
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
GPT-6.1 Sol is here.
Upgraded with stronger agentic coding and computer use, near-Astra performance, and cached input at a 95% discount to standard input pricing.
GPT-6.1 Sol is built for complex refactors, deep codebase investigations, and long-running agents across apps.
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
Iβll explain the new Pro 200 plan differently, before I start live tweeting from DevDay on things that are going out!
Today we are going to ship a number of things that increase what you can do across the Plus and Pro plans. A lot of compute is online for this increase. As we increase the floor, we are changing the relative difference between plans to be
Plus = 1X
Pro 100 = 5X
Pro 200 = 10X
and we are reopening subscriptions for Pro 200 (we had paused it). If you have an existing plan you will keep the 20X multiplier for a bit and also receive a lot of additional credits because we know changes are hard even if it means that everyone will get more in the end.
Qwen3.8-Flash-Next - 4.1vπ₯³
π Peak c=1 @ 118 tok/s
π Average c=1 @ 80 tok/s
π Peak c=64 @ 834 tok/s
βοΈ Prefill 3,233 tok/s
π» TTFT 0.42s
β‘οΈ github.com/myllmbox/qwen38-fβ¦
Agree, FB reinvented OpenClaw for normal people. Gave it a furry ass and called it a day. And there is a thing: normal people do not ask these kinds of questions - hey, can I put my model in it? They use. Sharing their data with FB, getting more and more locked inside the ecosystem.
That is exactly what Mark wants.
Tbh, every major player on the market wants the same.
As a result, you don't own your data or your life anymore. And the word "freedom" sounds very different these days.
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
New acceleration method for Minimax H3: Veda Sparse
huggingface.co/Veda-Sparse/Mβ¦
Oh, an Anthropic "Harry Potter" was a PR experiment all the way. Who can predict that? (irony)
Former Anthropic researcher Jacob Coxon became a media superstar after going public with his AI fears. He insisted that he wasnβt working with any third party organizations. Familiar sources told us a different story: DEY., a PR firm representing many of the most prominent AI safetyists, was booking his interviews.
One source, who had direct knowledge, even said DEY. preemptively booked Nate Soares, a prominent AI safety figure, for interviews that directly overlapped with Jacob going public.
Jacob working with DEY. is notable for two reasons: first, as mentioned, he previously said he wasnβt working with third parties. Second, we are in the middle of a national conversation about the future of AI that is actively determining how we regulate the most powerful technology in the world, largely thanks to the panic stirred up by Jacob β and itβs in the publicβs interest to know who, exactly, is behind it.
Scoop from @huntryerson π
o1-preview just 2 years ago? omg, it felt like an eternity
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to βthinkβ before answering
Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
Introducing Julia-1:
Our first classification model that runs on almost anything.
Learn more π
supersoniclabs.ia.br/julia-1β¦
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
π¨ OpenAI Set to Reveal a Long-Term Agent at DevDay
- Codenamed "Aeon" β built for long-running tasks, similar to Grok Bot or Manus, working for hours, days, even weeks
- Likely built on Astra, already strong at long-horizon work
- Runs in a cloud environment like Cursor β sets everything up remotely and keeps grinding until the task's done
- OpenAI already has the infra (hosted sandboxes, multi-agent workflows) to make this the natural next step
- Rumored to support multiple agents collaborating on the same task β if real, that's a big deal
Slava π¨π¦ β€οΈ πΊπ¦ retweeted
Weβre launching the Army of Robots.
In 2022, we launched the Army of Drones. Today, drones account for over 95% of battlefield strikes, and Ukraine has more than 700 UAV manufacturers.
Now we need the next technological breakthrough: the robotization of warfare.
The goal is simple β save lives. Robots should take on the most dangerous missions: evacuating the wounded, delivering ammunition, mining and demining, reconnaissance, defending positions and engaging targets.
The Army of Robots is not one company. Itβs an ecosystem. We will invest in defense tech companies, launch our own technology projects, test them with the military and scale what works on the battlefield.
Weβre now looking for defense tech companies and engineers working on robotic technologies β as well as a CTO / Tech Lead for the Army of Robots.
Join us: thearmyofrobots.com/en
Qwen3.8-Flash-Next on solo GB10 just got better.
Totally new optimized PLE offload. Allows to get ~20% more speed π and we again have lots of KVβοΈ
c=1 -> 82 token/s βοΈ
c=16 -> 318 token/sβοΈ
β‘οΈgithub.com/bilikaz/qwen38-flβ¦
Qwen3.8-Flash-Next @ sustainable ~50β51 tok/s for code π€―
While for thinking and code it gets ~42 tok/s π
v1 had good moments. v2 has speed, and it keeps it.
β‘github.com/bilikaz/qwen38-flβ¦
RTX PRO 6000 96GB + DGX Spark owners rejoice! π₯
You can now run the highest quality Xiaomiβs MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process π
309B total parameters. Only ~15B active per token. Xiaomiβs published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5.
And on ONE RTX PRO 6000:
184.1 tok/s p50 with DFlash
49.6 tok/s without drafting
2,258+ tok/s prefill
321.9 tok/s confirmed aggregate @ C=8
π£πππ π¬π’π¨π₯ πππ₯π
RTX PRO 6000 96GB:
2.20 bpw SAGE-EXL3
86.94 GB
DGX Spark 128GB:
2.50 bpw SAGE-EXL3
98.48 GB
The recipe includes separate profiles for both systems. The 184.1 tok/s result is specifically from the RTX PRO 6000; Spark DFlash gains are currently lower and more prompt-dependent.
πππππππ§π¬ π₯πππππ£π§π¦
Across two completely untouched holdout sets:
2.20 bpw:
82.16β82.64% top-1 agreement
0.1926β0.1986 mean KLD
2.50 bpw:
83.76β87.30% top-1 agreement
0.1086β0.2055 mean KLD
Top-1 agreement means the EXL3 pack selected the SAME most-likely next token as Xiaomiβs reference.
KLD compares the entire next-token probability distribution. Lower means the quant tracks the reference more closely.
ππ‘ ππ π£π’π₯π§ππ‘π§ πππ§πππ
Xiaomi does not publish a full BF16 expert checkpoint.
The 303B routed-expert bank already ships in MXFP4. Attention ships in block-FP8, with embeddings, norms and the output head in BF16.
So these fidelity numbers compare against Xiaomiβs actual released mixed-precision checkpoint running through its official implementation, not against a hidden full-BF16 teacher.
This is effectively EXL3 agreement with the best public source that exists.
ππ’πͺ πππ’π¦π ππ¦ ππππ¦π π§π’ π§ππ ππ₯π’π‘π§πππ₯?
Xiaomiβs published numbers:
AutomationBench:
MiMo Flash 52.3
Claude Opus 5 50.3
GPT-5.6 Sol 45.8
Terminal-Bench 2.1:
MiMo Flash 87.6
Claude Opus 5 89.1
GPT-5.6 Sol 88.8
OSWorld:
MiMo Flash 80.8
Claude Opus 5 83.4
GPT-5.6 Sol 83.0
VisualCoding:
MiMo Flash 71.5
Claude Opus 5 70.0
GPT-5.6 Sol 73.4
Artificial Analysis has not scored Flash independently yet.
The larger MiMo-V2.6-Pro sibling scores 46 on the AA Intelligence Index, tied with Grok 4.7 and sitting beside GLM-5.3 at 45.
This release is the text backbone. The original DFlash speculative drafter is repaired and wired through the recipe; vision and audio towers are not included yet.
Huge credit to @XiaomiMiMo team for the model and to @turboderp_ / ExLlamaV3 for the EXL3 engine and format. @XiaomiMiMoDevs
EXL3 weights + fidelity results:
huggingface.co/vcruz305/MiMoβ¦
Complete serving recipe:
github.com/vcruz305/MiMo-V2.β¦