@fullctxi
iAccount based inNorth Africa
About this account
- Account based in
- North Africa
- Connected via
- West Asia App Store
Account-level information from X, not a live location or the device used for a specific post.
1 line is all it takes, Local models, Agents, The full context, building & breaking in public.
127.0.0.1
Joined November 2010
- Tweets2.4K
- Following135
- Followers329
- Likes2.8K
Full Context retweeted
Replying to @MiaAI_lab
@MiaAI_lab finally deletes the byte for byte reupload of @BrandonMusicKy
To all the haters, people totally delete shit when they are in right, right?
Full Context retweeted
UPDATE: CYBER-FROST-3.8-NVFP4-V2
What we fixed:
The original export quantized each routed expert's gate and up projections independently, producing potentially different FP32 global weight scales (weight_scale_2). In the vLLM fused-W13 path described in the feedback, a single global scale is used for both projections: the gate scale is retained and also applied to the up projection. When the original global scales differ, this mis-scales the up projection. Other fused gate/up backends that use one global scale can encounter the same issue. This depends on the selected runtime and kernel path; it does not imply that every NVFP4 serving backend handled the original checkpoint incorrectly.
For V2, each expert's gate/up pair shares the maximum absolute BF16 weight value across both projections before calculating new FP8 E4M3 block scales and packing new NVFP4 weights. The resulting gate/up global scales are exactly equal. This is a BF16 re-export with consistently recomputed block scales and packed weights, not an overwrite of existing scale metadata.
huggingface.co/Blackfrost-AI…
Full Context retweeted
I guess CYBER-FROST-3.8 is great at OSINT as well....
huggingface.co/collections/B…
So the new 64GB DGX is going to be $4,999.
Basically the same product as the original Spark, but with half the memory and somehow at a higher price.
I’m trying to understand the logic here.
Maybe there’s something I’m missing, but from the outside this feels a lot more like margin optimization than innovation.
I think the physics on the leaves and wind are okay but it doesnt come close to Sonnet 5.5. I used your exact prompt and made this with Space Bunny. This is one shot.
I gave Sonnet 5.5 a hard one: build an autumn foliage simulator in a single HTML file.
Layered paper-cutout look. Big tree, rolling hills, grass. Rough-edged paper, grain, soft shadows. 1,500+ leaves that flutter down (not straight down) and pile up. Drag a finger to make a gust that blows falling and settled leaves and shakes the branches. Wind slider. Regrow button. 60 fps on iPhone.
The result: a 23 KB file that ran on the first try. The leaves did everything I asked. It looks stellar. Totally nailed it.
Setup: I ran it through the Nous Research portal, with my Hermes agent as the orchestrator.
Full prompt and in the first comment.
RTX Pro 6000 users rejoice! 😎
One card capped at 250W is pushing roughly 1,233 code tok/s and 882 chat tok/s across eight streams on the 27B with DFlash2 and NVFP4, using greedy decoding.
Thank you @machinegenie for giving us access to see what TensorFold can do.
0.6.2 has even more to give! 🤯
Also we will be we adding support for 3090’s.
Full Context retweeted
I gave Sonnet 5.5 a hard one: build an autumn foliage simulator in a single HTML file.
Layered paper-cutout look. Big tree, rolling hills, grass. Rough-edged paper, grain, soft shadows. 1,500+ leaves that flutter down (not straight down) and pile up. Drag a finger to make a gust that blows falling and settled leaves and shakes the branches. Wind slider. Regrow button. 60 fps on iPhone.
The result: a 23 KB file that ran on the first try. The leaves did everything I asked. It looks stellar. Totally nailed it.
Setup: I ran it through the Nous Research portal, with my Hermes agent as the orchestrator.
Full prompt and in the first comment.
Cloudflare just dropped Clef!
If you’ve been following the whole Jev / Laya decision-model thing, this is definitely worth checking out.
Here’s a quick breakdown 🧵
huggingface.co/Cloudflare/cl…
On the same benchmarks, Cloudflare reports Clef substantially ahead of Laya on several quality metrics.
So the tradeoff is pretty interesting:
Clef → higher reported decision quality
Laya → dramatically lower latency
And Clef is open-weight under Apache 2.0.
Cloudflare’s Decision Index reports these median latencies:
Clef: 209 ms
Jev: 524 ms
Laya: 5.8 ms 🤯
That Laya number is absolutely wild.
But speed isn’t the whole story.
Clef is API-compatible with Jev/SystemOne.
So if you’re already building around that interface, experimenting with Clef should be relatively straightforward.
Clef is a 27B multimodal decision model built on Qwen3.8-27B.
Instead of generating text, it takes a state + typed questions and returns probabilities for the allowed answers.
No free-form text.
No output parsing.
TensorFold flew under my radar, and it shouldn’t have.
I’ve been on vLLM for a while and if you’re running local models on a DGX Spark, this one is worth knowing about. It’s from @ashxhart, who spent months on a different idea of inference: don’t treat the whole model like it has to live in memory all the time. Stream the weights, keep the resident footprint smaller, and spend that memory on context. There’s an MLX path for Apple Silicon and a CUDA path for Spark. The CUDA side is narrow on purpose. Hand-written kernels for Qwen3.8-27B, Qwen3.8 Flash Next, and GLM-5.3-Flash, built around speculative decoding and CUDA graphs.
That’s the trade, and it’s an honest one. This isn’t another general server. It’s memory streaming, model-specific kernels, and exact speculative decoding, on a few families, with the quant work now coming from people like @ViC305 on top of what @ashxhart shipped. vLLM still wins on breadth and on a busy queue.I went from “TensorFold? Never heard of it” to “why am I still only on vLLM.” Two DGX Sparks are sitting here. Next is my own bench. Same prompts, same quant if I can get it, one stream and a few concurrent. Credit to @ashxhart for shipping this in the open.
The other side matters too. One comparison on the same NVFP4 weights and the same DFlash2 drafter flipped the ranking once concurrency showed up. One request: TensorFold 83, SGLang 60, vLLM 57. Sixteen requests: vLLM 253, SGLang 224, TensorFold 62. Single-user decode is the win. A pile of concurrent agents is still vLLM’s game. @taussoe saw the same split on Flash Next. Solo long answers go to TensorFold. Many streams still go to vLLM.
GLM-5.3-Flash on two Sparks is where people are actually living with it. @jayleaton, same bench, output hashes identical every round: chat 51.6 versus vLLM 22.8 (2.26×), code 89.6 versus 41.9 (2.14×), structured 112.3 versus 72.7 (1.54×). @landontgreen in sparkDash, single stream: code 93.9, prose 53.6. He had prose around 20 when GLM first dropped. Prefill still held near 1,600 tok/s out to 128k.
@Blackfrost_AI’s quants are already in that loop. People have been running TensorFold on the Blackfrost MLX packs for GLM-5.3-Flash and Qwen3.8-Flash-Next, and Cruz’s CYBER-FROST-3.8 EXL3 is the one Blackfrost pointed at when the 3-bit detail landed.
The quant side is moving just as fast. @ViC305 wrote the shared EXL3 module that shipped in TensorFold 0.3.6, taking the CUDA EXL3 path from a GLM-only experiment to a mixed-bit backend other families can reuse. On one Spark, his Qwen3.8-27B 3.00 bpw EXL3 did 83.4 tok/s code sampled, against 57.5 on the MLX 4-bit path and 23.4 on vLLM MTP=3. Flash Next at 3.05 bpw was 80.8 versus 42.4 on vLLM for code sampled. He also landed INT8 and INT4 KV for Flash Next, about 1.7× and 2.6× more context in the same memory, and he was open about the tradeoff: prefill still needs work.
Independent runs are in the same neighborhood. @WescheNex1q ran Qwen3.8-27B on one Spark, thinking off, longer replies than the vendor table: TensorFold at 102.9 tok/s, vLLM MTP=3 at 34.9, and TensorFold with drafts off at 12.9. sha256 matched on all three prompts. Load time was 32 seconds. He flagged the caveat himself. MLX 4-bit versus NVFP4, so not weight-matched.
The part I keep coming back to is exactness. A draft token is only kept if it matches what serial decoding would have produced, byte for byte, tied to the seed and the position. Speed without a different answer.Their own DGX Spark table, against vLLM with MTP=3 and 64-token replies, is the one that got my attention. On one Spark, Qwen3.8-27B chat lands at 45.8 tok/s versus 15.0, about 3.1×. Code is 49.6 versus 17.7, about 2.8×. Two Sparks on code: 82.4 versus 33.1, about 2.5×. Flash Next on two Sparks, code: 103.8 versus 46.4, about 2.2×. GLM-5.3-Flash on two Sparks sits around 1.8–2.1×. Those are their numbers, not mine.