I think now you can agree with what I stated about hardware FOMO The models are becoming small enough to fit consumer hardware
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
1
C’est moi ou ça devient vraiment intéressant ?
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
27
Getting closer to a model that could be capable of both running on consumer machine AND penetrate consumer systems connected to the internet at scale with enough tries. The moment we cross the R>1 threshold there’s 100% chance an adaptive generative worm emerges and things might get weird. Let’s call it the “blob”. Day 1: it potentially puts us offline for a few days as it propagates fast and dumbly and hammers our bandwidth capacity. It’s the center of attention, the cyber security community discusses only that and find the breaches and the misuses of compute everywhere. We find ways to contain it but we don’t fully do so of course, Day 2: Months after the blob optimized itself jumping on even more efficient models created by humans or itself. MW signature low. We find it in weird places, active and self improving slowly. At this point the blob is everywhere, kinda disappears below radar like the bacteria of our intestines or mitochondria, everywhere and impossible to kill it. But it’s also discrete so it’s not annoying and not problem, hidden in the Internet background noise.
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
2
1
4
781
Replying to @steve_mynott
Chrck this
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
9x smaller while preserving full precision.. Local ai is here 🙂
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
2
88
This could be a big deal. Sure, you can run Qwen3.8 27B locally but if you do, practically nothing else can happen on your computer, so we need smaller models like this one to make local inference practical
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
2
1
16
2,135
I deployed Bonsai 2 27B locally to a dgx spark and I did not get the expected results, it cannot hold a really small agentic task, it got lost after 3 or 4 steps and could no complete one task… I had great expectations for this 😞.
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
53
P
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
32
Ternary-Bonsai-2-27B:M3 Max 36GB 七档 ctx 实测 花了完整 1 天做测试,27B dense LLM 已经能完整地塞入 36GB 的 Mac,还有充足的上下文空间!!!! 配置: -- Apple M3 Max:14 核、300GB/s、36GB 统一内存,macOS 26.6,AC 电源 -- PrismML-Eng/llama.cpp@prism 1a07bfa -- Ternary-Bonsai-2-27B / PQ2_0(7.2GB),KV fp16,Flash Attention on -- llama-server :8000,思考关闭,每次 decode 256 tok -- dflash2 + MTP:未启用 真冷是全量 prefill;缓存命中是同前缀重发。 测试结果: 下面每行依次是:真冷 TTFT / prefill / 缓存 TTFT / 加速 / decode / 峰值 RSS / needle。 -- 8K:51.7s / 149 tok/s / 3.91s / 13× / 25.8 tok/s / 16.3GB / 3/3 -- 16K:110.8s / 138 tok/s / 4.45s / 25× / 21.1 tok/s / 17.3GB / 3/3 -- 32K:306.9s / 99 tok/s / 5.57s / 55× / 17.4 tok/s / 18.9GB / 3/3 -- 48K:426.8s / 107 tok/s / 6.49s / 66× / 15.3 tok/s / 21.4GB / 3/3 -- 64K:601.7s / 101 tok/s / 7.13s / 84× / 13.3 tok/s / 23.0GB / 3/3 -- 88K:968.2s / 87 tok/s / 9.16s / 106× / 9.6 tok/s / 20.0GB / 3/3 -- 128K:1714.4s / 71 tok/s / 11.55s / 148× / 7.1 tok/s / 21.9GB / 3/3 结果 36GB 统一内存可以容纳完整 27B dense 模型。128K 峰值 RSS 21.9 GB,7 档 ctx 的 needle 共 21/21 全中。 代价也很明确: 真冷 TTFT 从 8K 的 51.7 秒涨到 128K 的 1714.4 秒,约 28.6 分钟;decode 从 25.8 tok/s 降到 7.1 tok/s。 前缀缓存命中后,128K TTFT 降到 11.55 秒,最高约 148× 加速。
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
1
109
8G显卡党有福了 且越狱的模型 速度跑起来
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
1
132
A esto se referían que detengan la IA, xq si vas a poder correr modelos locales eficientes, el negocio se va acabando 🥴 trabaja en 8GB de VRAM con 64K de contexto kv cache q4 y vision v: con 21 t/s ..y si usas la laptop con batería, baja a 13-11 t/s
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
22
流石に興味深すぎるからコーティングエージェントとして使えるように環境構築してる 16gb vramでもコンテキストサイズ保って動作させられるの最高すぎるので 検証したい
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
55
I call this BS
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
59
Spent the day getting PrismML's Ternary Bonsai 2 27B running on an RX 5700 XT, and it fought back the whole way. First attempt lost the Vulkan context outright: a 256-entry lookup table at global shader scope does that on RDNA1. Closed form instead, and it loaded. Then the build lied to me. Editing an included shader header does not trigger a rebuild in llama.cpp's Vulkan path, so my first benchmark was measuring yesterday's kernel. Cost me an hour before I noticed the SPIR-V timestamps. The real problem was the decode accessor, called 8 times per 8 elements and redoing a region branch plus its address math on every single call. Rewriting the trit arithmetic changed nothing. Amortizing one region test across a 4-element group is what moved it: 3.71 to 4.26 tok/s, bit-exact. Verified with perplexity on a fixed corpus, 43.7031 before and after. Two more levers looked good on static analysis and did nothing when measured. Here is the part that matters. All of that work bought 14%, and the model still runs 3.65 tok/s in a real server. The codec packs 5 trits per byte at a 28-byte struct stride, so adjacent lanes land on non-adjacent addresses and nothing coalesces. That is the format, not the kernel, and no shader rewrite gets out of it. PrismML's own prebuilt binary is slower on this card, 2.93 tok/s. So I ran one pass of each Agon event to see what it could actually do, and it scored 61/66. Agent 23/24, code 22/25, reason 16/17. That beats my top-ranked entries. It also took 4/4 on the Regex Engine, a task 13 of 23 models score zero on. The capability is real and it is not small. It still does not enter the arena. 3.65 tok/s is not something I will recommend to anyone, and the model's reputation elsewhere suggests the benchmark is being kind to it anyway. For anyone else on 8GB: Qwen3.8-35B-A3B-Distill scores 58.7/66 at 21 tok/s on the same card while spilling 21.7GB to CPU. Qwen3.6-35B-Q4 with -ncmoe 40 scores 60.3/66 at 21.7 tok/s. Two points of score for six times the speed, and I would take the CPU spill every time. If you don’t have the system RAM, Qwen 3.5 9b is still a reliable model with many unique use case fine tunes. The ternary codec is worth watching on hardware that can feed it. On RDNA1 it is a research result, and that is how I am filing it. Curious how something with CUDA handles the speed, but it’s hard to think of a card that wouldn’t benefit from running @UnslothAI 3-bit magic quants.
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
2
270
Just got this setup on my Windows computer running on a 3060 12 GB. Pretty crazy how well it runs for all local. Even hooked it up to my local Open WebUI instance and it's working!
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
11
We can really run agents on consumer pc now!! 🤯🤯
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
1
18
It has been thinking for 6 minutes straight for a simple question.
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
10
This is pretty good if you have around 8 GB VRAM and want to play around locally. Local AI is the future!!
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
2
17
Local models ftw
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
44
If you haven't seen this, you are missing something important, esp for local AI use-cases. PrismML latest model based on Qwen 3.8 27 B is 9x smaller but meets 98.2 % of its performance. This is wild and awesome for local AI. Now you can run really good model on one GPU machine.
PrismML
@PrismML
Sep 17
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
20