@pmbstuff

Full Stack AI Engineer & Researcher. Building AI-powered tools.

Canada
Joined July 2010
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
GPT-6.1 Sol is here. Upgraded with stronger agentic coding and computer use, near-Astra performance, and cached input at a 95% discount to standard input pricing. GPT-6.1 Sol is built for complex refactors, deep codebase investigations, and long-running agents across apps.
151
243
73
3,850
142,845
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
Introducing dots, powered by GPT-6 Astra. Remarkably capable, always-on agents built to handle everything.
1,145
1,694
1,652
21,190
4,064,586
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today.
433
895
600
12,990
1,149,029
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
I’ll explain the new Pro 200 plan differently, before I start live tweeting from DevDay on things that are going out! Today we are going to ship a number of things that increase what you can do across the Plus and Pro plans. A lot of compute is online for this increase. As we increase the floor, we are changing the relative difference between plans to be Plus = 1X Pro 100 = 5X Pro 200 = 10X and we are reopening subscriptions for Pro 200 (we had paused it). If you have an existing plan you will keep the 20X multiplier for a bit and also receive a lot of additional credits because we know changes are hard even if it means that everyone will get more in the end.
3,037
436
1,053
9,088
1,620,390
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
"Come play with me"
52
171
31
2,525
63,685
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
Qwen3.8-Flash-Next - 4.1vπŸ₯³ πŸš€ Peak c=1 @ 118 tok/s πŸ“Š Average c=1 @ 80 tok/s 🏭 Peak c=64 @ 834 tok/s βš™οΈ Prefill 3,233 tok/s πŸ’» TTFT 0.42s ⚑️ github.com/myllmbox/qwen38-f…
7
9
91
7,905
Expected
Deleted Muse after seeing this post on Threads about how it told some Facebook Marketplace sellers the guy’s address and they showed up at his door Dangerous and creepy This would have been 1000x worse if the person was a woman
6
Agree, FB reinvented OpenClaw for normal people. Gave it a furry ass and called it a day. And there is a thing: normal people do not ask these kinds of questions - hey, can I put my model in it? They use. Sharing their data with FB, getting more and more locked inside the ecosystem. That is exactly what Mark wants. Tbh, every major player on the market wants the same. As a result, you don't own your data or your life anymore. And the word "freedom" sounds very different these days.
I think I speak for many when I say, the people want freedom of choice. You've built a great harness, computer-use agent, and supporting backend - but people don't want model lock-in. Grokbot cockblocked itself the same way and it's just not as powerful as running something like Astra/Fable/Opus.
27
I smell fear. Muse looks like a crab's butthole, btw.
15
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
New acceleration method for Minimax H3: Veda Sparse huggingface.co/Veda-Sparse/M…
9
11
1
197
13,428
Yes, more of that please!
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5. We're launching NerfBench on BridgeBench tomorrow. We have the day 1 results. Tomorrow morning we show you the retest. Is Claude Opus 5.5 nerfed or not?
3
Oh, an Anthropic "Harry Potter" was a PR experiment all the way. Who can predict that? (irony)
Former Anthropic researcher Jacob Coxon became a media superstar after going public with his AI fears. He insisted that he wasn’t working with any third party organizations. Familiar sources told us a different story: DEY., a PR firm representing many of the most prominent AI safetyists, was booking his interviews. One source, who had direct knowledge, even said DEY. preemptively booked Nate Soares, a prominent AI safety figure, for interviews that directly overlapped with Jacob going public. Jacob working with DEY. is notable for two reasons: first, as mentioned, he previously said he wasn’t working with third parties. Second, we are in the middle of a national conversation about the future of AI that is actively determining how we regulate the most powerful technology in the world, largely thanks to the panic stirred up by Jacob β€” and it’s in the public’s interest to know who, exactly, is behind it. Scoop from @huntryerson πŸ‘‡
1
5
o1-preview just 2 years ago? omg, it felt like an eternity
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to β€˜think’ before answering Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.
6
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
Introducing Julia-1: Our first classification model that runs on almost anything. Learn more πŸ‘‡ supersoniclabs.ia.br/julia-1…
108
274
97
3,228
293,305
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
🚨 OpenAI Set to Reveal a Long-Term Agent at DevDay - Codenamed "Aeon" β€” built for long-running tasks, similar to Grok Bot or Manus, working for hours, days, even weeks - Likely built on Astra, already strong at long-horizon work - Runs in a cloud environment like Cursor β€” sets everything up remotely and keeps grinding until the task's done - OpenAI already has the infra (hosted sandboxes, multi-agent workflows) to make this the natural next step - Rumored to support multiple agents collaborating on the same task β€” if real, that's a big deal
76
156
75
2,214
240,693
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
We’re launching the Army of Robots. In 2022, we launched the Army of Drones. Today, drones account for over 95% of battlefield strikes, and Ukraine has more than 700 UAV manufacturers. Now we need the next technological breakthrough: the robotization of warfare. The goal is simple β€” save lives. Robots should take on the most dangerous missions: evacuating the wounded, delivering ammunition, mining and demining, reconnaissance, defending positions and engaging targets. The Army of Robots is not one company. It’s an ecosystem. We will invest in defense tech companies, launch our own technology projects, test them with the military and scale what works on the battlefield. We’re now looking for defense tech companies and engineers working on robotic technologies β€” as well as a CTO / Tech Lead for the Army of Robots. Join us: thearmyofrobots.com/en
891
1,787
790
11,109
2,914,598
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
Qwen3.8-Flash-Next on solo GB10 just got better. Totally new optimized PLE offload. Allows to get ~20% more speed πŸš€ and we again have lots of KV❗️ c=1 -> 82 token/s ❗️ c=16 -> 318 token/s❗️ ⚑️github.com/bilikaz/qwen38-fl…
Qwen3.8-Flash-Next @ sustainable ~50–51 tok/s for code 🀯 While for thinking and code it gets ~42 tok/s πŸš€ v1 had good moments. v2 has speed, and it keeps it. ⚑github.com/bilikaz/qwen38-fl…
19
16
1
198
25,926
Slava πŸ‡¨πŸ‡¦ ❀️ πŸ‡ΊπŸ‡¦ retweeted
RTX PRO 6000 96GB + DGX Spark owners rejoice! πŸ”₯ You can now run the highest quality Xiaomi’s MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process πŸš€ 309B total parameters. Only ~15B active per token. Xiaomi’s published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5. And on ONE RTX PRO 6000: 184.1 tok/s p50 with DFlash 49.6 tok/s without drafting 2,258+ tok/s prefill 321.9 tok/s confirmed aggregate @ C=8 π—£π—œπ—–π—ž 𝗬𝗒𝗨π—₯ 𝗖𝗔π—₯𝗗 RTX PRO 6000 96GB: 2.20 bpw SAGE-EXL3 86.94 GB DGX Spark 128GB: 2.50 bpw SAGE-EXL3 98.48 GB The recipe includes separate profiles for both systems. The 184.1 tok/s result is specifically from the RTX PRO 6000; Spark DFlash gains are currently lower and more prompt-dependent. π—™π—œπ——π—˜π—Ÿπ—œπ—§π—¬ π—₯π—˜π—–π—˜π—œπ—£π—§π—¦ Across two completely untouched holdout sets: 2.20 bpw: 82.16–82.64% top-1 agreement 0.1926–0.1986 mean KLD 2.50 bpw: 83.76–87.30% top-1 agreement 0.1086–0.2055 mean KLD Top-1 agreement means the EXL3 pack selected the SAME most-likely next token as Xiaomi’s reference. KLD compares the entire next-token probability distribution. Lower means the quant tracks the reference more closely. 𝗔𝗑 π—œπ— π—£π—’π—₯𝗧𝗔𝗑𝗧 π——π—˜π—§π—”π—œπ—Ÿ Xiaomi does not publish a full BF16 expert checkpoint. The 303B routed-expert bank already ships in MXFP4. Attention ships in block-FP8, with embeddings, norms and the output head in BF16. So these fidelity numbers compare against Xiaomi’s actual released mixed-precision checkpoint running through its official implementation, not against a hidden full-BF16 teacher. This is effectively EXL3 agreement with the best public source that exists. 𝗛𝗒π—ͺ π—–π—Ÿπ—’π—¦π—˜ π—œπ—¦ π—™π—Ÿπ—”π—¦π—› 𝗧𝗒 π—§π—›π—˜ 𝗙π—₯π—’π—‘π—§π—œπ—˜π—₯? Xiaomi’s published numbers: AutomationBench: MiMo Flash 52.3 Claude Opus 5 50.3 GPT-5.6 Sol 45.8 Terminal-Bench 2.1: MiMo Flash 87.6 Claude Opus 5 89.1 GPT-5.6 Sol 88.8 OSWorld: MiMo Flash 80.8 Claude Opus 5 83.4 GPT-5.6 Sol 83.0 VisualCoding: MiMo Flash 71.5 Claude Opus 5 70.0 GPT-5.6 Sol 73.4 Artificial Analysis has not scored Flash independently yet. The larger MiMo-V2.6-Pro sibling scores 46 on the AA Intelligence Index, tied with Grok 4.7 and sitting beside GLM-5.3 at 45. This release is the text backbone. The original DFlash speculative drafter is repaired and wired through the recipe; vision and audio towers are not included yet. Huge credit to @XiaomiMiMo team for the model and to @turboderp_ / ExLlamaV3 for the EXL3 engine and format. @XiaomiMiMoDevs EXL3 weights + fidelity results: huggingface.co/vcruz305/MiMo… Complete serving recipe: github.com/vcruz305/MiMo-V2.…
5
10
1
72
5,553