@ali7ep

Co-Founder of https://nitter.cf/t.co/zvPI9BnOaK

Joined July 2022
Ali 🦁 retweeted
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
153
451
228
4,059
604,615
Ali 🦁 retweeted
astra patched a few bytes in my Windows XP ISO to avoid a memory check that crashes the game this is incredible
29
16
931
52,338
Ali 🦁 retweeted
Bonsai 2 27B dropped on Thursday. PrismML reports 98.2% aggregate performance retention versus Qwen3.8-27B FP16: 83.9 vs. 85.4 across a 20-benchmark suite. That is an improvement from the 95% aggregate retention reported for the previous Bonsai 27B generation (note the underlying benchmark suite changed). The accompanying whitepaper provides substantially more detail on individual benchmarks and evaluation methodology - this makes the results easier to examine beyond the headline average. Looking through the 20 benchmarks in granularity, performance retention varies meaningfully by task. Some of the larger drops include: OCR Bench v2: 93.3% retention BigCodeBench: 94.4% GPQA Diamond: 94.8% At the same time, a few benchmarks like IFBench and HumanEval scored above 100% which partially offset the underperformance to get to the 98.2% average. One important, overlooked part of the whitepaper is that PrismML also separately reports two longer-horizon agent evaluations: Terminal-Bench 2.1 Bonsai 2: 52.8 FP16: 69.7 75.8% retention SWE-bench Verified Bonsai 2: 60.8 FP16: 80.6 75.4% retention These two benchmarks are not included in the 20-benchmark suite used for the 98.2% headline. The paper reports them separately and does not explain why they were excluded from that aggregate. What sets these two benchmarks apart is that they test long-horizon, multi-step engineering and terminal-agent tasks where the model repeatedly reasons, acts, observes results, updates state, and continues over many turns. That matters at extreme compression levels because small model errors can accumulate across long generation trajectories. PrismML itself notes that these workloads are particularly sensitive to model degradation because small errors can compound over many steps. That makes the ~75% retention on these two benchmarks an important data point for anyone considering heavily compressed models for running long-running agents. One other metric that would be useful to see in future releases is output-token efficiency. It has been widely reported and observed that low-bit reasoning models can increase reasoning length, retries and semantic repetition. This makes tokens-per-success and reasoning-token counts useful metrics to report alongside accuracy. Overall, Bonsai 2 is a good progress in extreme compression. For those looking to deploy extremely quantized models for real world tasks, please evaluate the performance carefully so you are getting the most out of the model for your specific use case and needs.
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
6
5
3
71
8,156
There was a time that I joined a company for the money and the rep of the CEO. I hated the product and never used it (personal financial management app) but as a TechLead I jumped into issues and fixed them, I just had to have a slight understanding of what was going on.
I am done with this shit. It is over. The state of engineering right now is horrible. It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking. I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens.
1
26
No deep understanding of the product was needed... People working there had super better knowledge of how the product was working and how the things had to be, but they weren't experts at steering the development and fixing the bugs.
1
4
I was somehow the CC of the team and they kept pushing Enter.
2
Ali 🦁 retweeted
Jev is great at zero-shot classification, but specialist classifiers will dominate commercial use cases. @trycua tuned a tiny model that scored 99.7% on their form-filling eval. Hosted Jev scored 83.6%. I tuned GLiNER 2.5 on a task in 51 minutes yesterday and it crushes Jev. And it's local. And 8.8x faster: You too can do this. Linked post in comments.
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
31
38
12
434
69,551
Ali 🦁 retweeted
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
115
478
154
5,396
1,208,624
Ali 🦁 retweeted
JevBench results are in. Jev still in the lead, but it's close.
Introducing JevBench. The first benchmark for Jev class models. Original Jev by @typesafeai in the lead at 75.3. SemIf #2 at 74.6. All results at benchmarkheaven.com/jev-mode…
53
90
35
937
173,749
98.2% of Qwen3.8 27B my ass. I got hyped and gave it a real agent job right away. Build me an FPS in three.js, 6 hours on a 3090. My most standard and default prompt that I always use. It spent the first 32K tokens on a plan without writing a single file, then shipped a black screen, 2 shaders that dont compile and a player who spawns dead. And wrote "verified" in the final report. Ok, too hard. I gave it the easiest thing I have, a voxel pagoda garden in one html file. Video attached. 3 hours for THIS. On the way it deleted its own file and spent an hour debugging a raycaster nobody asked for. Where the 98.2% is in all this I have no fucking idea. Same Qwen3.8 27B, same pagoda task, ISTA-DASLab GSQ-RCO IQ2_XS on a 12 GB 3080 Ti gave me day and night, real shadows and koi fish in the pond. 8.4 GB on disk, real 2.50 bpw, 131072 ctx with q4_0 KV, 47 tok/s at 128K. 2.5 bits beats 2.13 bits by a lot when the 2.13 is this shit. Post with the video and the exact command in the replies. In the replies, I'll attach what a proper Qwen3.8 27B created in my hands. Bottom line: if you have at least 12GB VRAM, use ISTA-DASLab GSQ-RCO of all THIS.
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
189
91
32
2,154
342,785
Invest in your marriage, not your wedding. I've been married for 30+ years, and we got married in the church basement under the Bingo board. I can't tell you how many big weddings I've been to over the years that ended in big divorces.
Expensive weddings are one of the worst possible uses of money.
57
43
2
973
20,479
Ali 🦁 retweeted
JUST IN: Trump invites Tim Cook, Sam Altman, Jensen Huang, & other tech leaders to the White House state dinner for China’s Xi Jinping.
67
350
29
6,913
377,709
Ali 🦁 retweeted
All the grifters are completely wrong about Jev’s architecture so I decided I’d release an open-weight version. BUT training takes time, so while we all wait I decided I’d drop the sauce. archerhume.com/posts/jevs-ar…
49
212
35
2,245
378,316
Rust is software cancer, Ross.
Rust is eating the world and Copilot brought a fork 🦀 @github used Copilot to rewrite Copilot’s agent runtime in Rust. 800K+ lines, 128 PRs, shipped incrementally. Existing end-to-end tests ran against the new code at every step. github.blog/ai-and-ml/genera…
23
Let's see what Fable 5.1 can cook
3
Rust is eating the world and Copilot brought a fork 🦀 @github used Copilot to rewrite Copilot’s agent runtime in Rust. 800K+ lines, 128 PRs, shipped incrementally. Existing end-to-end tests ran against the new code at every step. github.blog/ai-and-ml/genera…
34
Replaced my classifiers with Jev for agentic electronics part search. So far so good 👍 💯 It feels faster and more accurate than Gemini or Qwen. @typesafeai
92
Ali 🦁 retweeted
I ran a detailed comparison between Fable 5.1, Astra, and SWE-2 in Devin to map out their strengths and weaknesses: SWE-2 – Strengths: It explores territory other models never touch, looking at problems with a much wider lens. It picks up on subtle cues with very little context. You can trust it with minimal prompting, like telling it to delete something without worrying it might wipe out unrelated code. It has great discipline. When evaluating a problem, it targets the main pain point right away instead of wasting time on trivia. Its "what not to do" advice is genuinely useful and keeps you from spinning your wheels in solved areas. Unlike Astra, it avoids arbitrary, nonsensical refusals. It feels much more honest about what it can and cannot do. – Weaknesses: It lags far behind its base model, Kimi K3, in UI sense, aesthetics, and design taste. It also tends to stumble on a great idea and then hesitate to actually ship it. Astra – Strengths: Pure tenacity once it identifies a bug. It simply doesn't stop until it solves the problem. In one test, Opus 5 gave up and told me to file a support ticket, while Astra kept digging until it fixed the issue completely. It shows real engineering creativity, picking tight research paths with concise, structurally sound solutions and solid recommendations. It also has a great sense of success probability and offers precise statistical qualification for tasks. – Weaknesses: It tends to plan around small sample sizes and thin evidence. It gets bogged down in minor details instead of looking at the big picture and overall impact. It also struggles to synthesize ideas across different domains, overspecializing until the scope becomes too narrow. Fable 5.1 – Strengths: The highest evidence accuracy by far. It reliably delivers on what it promises, unlike models that promise everything and underdeliver. It has strong design taste, clear recommendations, and a holistic grasp of the end goal. It weighs priorities accurately. It can also pivot your entire strategy convincingly without drifting off into irrelevant territory. – Weaknesses: Its answers are dense and exhausting to read. Once it settles on an idea, it's hard to steer it away. Its research skims the surface rather than digging up rare findings. It can suffer from tunnel vision, sticking strictly to technical aspects while ignoring costs or viability metrics unless prompted. Patterns and Behaviors – Negation: Fable and SWE-2 set broad boundaries. Astra uses rigid negative rules that abruptly flip decisions when triggered. – Stability: Fable and Astra are steady. SWE-2 can be inconsistent, sometimes generating an idea only to backtrack on it. – Length: Astra is the most concise. Fable is the longest. SWE-2 sits in the middle, though it has repetitive filler that needs trimming. – Main Risks: SWE-2 over-explores possibilities at the expense of verifying existing code. Astra loses the forest for the trees. Fable turns practical technical decisions into academic literature reviews. – Standout Traits: SWE-2 wins on thoroughness, Astra wins on grit, and Fable wins on raw reasoning and persuasion. Scores and Recommendations – Scores: Fable 5.1: 8.5/10. Astra: 7.5/10. SWE-2: 7/10. – Pick Fable or SWE-2 if you want strict prompt adherence, clean execution, or deep mechanical logic. – Pick Astra if you need precise verification evidence and relentless debugging. – Scope and context: SWE-2 and Fable are tied, while Astra falls slightly behind. This is an initial assessment based on tests in Devin, and I'll need more runs to map out the nuances completely.
27
12
3
160
14,244
For C/C++ and low-level stuff, Fable 5.1 is on a whole other level; Astra has nothing to say...!
Testing Fable 5.1 on my Windows Kernel Driver Project...
12
Testing Fable 5.1 on my Windows Kernel Driver Project...
1
26