@geldekii
iAccount based inTurkey
About this account
- Account based in
- Turkey
- Connected via
- Turkey App Store
Account-level information from X, not a live location or the device used for a specific post.
Building software by day. Running local models at night so nothing leaves the house. AI/Agent builder 🛠️🤖 Indie Hacker 🚀
Joined January 2023
- Tweets2.4K
- Following2.1K
- Followers482
- Likes8.9K
111.2 tok/s on 2× RTX 3090 🔥
Huihui’s abliterated Swift Qwen3.8-Flash-Next IQ3_XXS running locally on Strata.
12,914 tokens into a coding-agent tool call.
64GB RAM + CPU offload.
256K context configured, vision enabled.
#LocalAI #Qwen #Strata #RTX3090
🤖 Made with AI
Erman Eroglu retweeted
İnsan sınırlı kapasitesi gereği, aklının almadığı şeyler için “bunu tanrı düşünmüş olmalı” sonucuna varıp tanrıyı ve dinleri var etti.
Öyle ya başka nasıl, kuşlar uçabilir, taşların arasından su çıkabilir, kadınlar doğurabilir, güneş her sabah doğabilirdi?
Bilgimiz arttıkça tanrının hayattaki yeri de azaldı ama güven ve inanma ihtiyacı sadece evrildi. Okuduğumuza inandık, devletimize güvendik.
Şimdi ilk defa bizden çok daha üstün düşünen makinalar ile karşılaşıyoruz.
Aynı anda milyonlarcamıza cevap veriyor, milyonlarca sorun çözülüyor, olaylar arasındaki ilişkileri analiz ediyorlar. Ve bunda her gün daha da iyi oluyorlar.
Hemen her dediğine inanmaya, yemeği önerdiği gibi yapmaya, önerdiği vitaminleri almaya razıyız.
Sonunda kendi yarattığımız tanrılar ile tanışıyoruz. Üstelik bu sefer varlar ve dualarımıza cevap da veriyorlar.
Gelecekte, eldeki inançların da yapay zekaya inanca ve güvene evrilmesi kaçınılmaz.
Ona “tanrı” demesek de inanıp güveneceğiz.
(bahçe ile uğraşırken kendimi felsefeye verdiğim doğrudur)
After this discussion just have grabbed my 3rd RTX 3090 for 900$. Lets get 4th and a threaddripper + wrx 80-90 if I can or a motherboard which supports 4 3090s at the same time. Lets see
🪶🚲 The scene is cute. The rig behind it is the interesting part.
A full Three.js beach world — Gerstner wave shaders, IK-rigged pelican legs, a seaside village, day↔night mode — built and revised 6 times in one continuous session by qwen3.8-flash-next-iq3_s, running on:
• 2× RTX 3090 (24 GB VRAM each) • 64 GB system RAM • 70–103 tokens/sec sustained generation • 256K context — the entire ~1,000-line codebase stayed in working memory across every iteration: rescaling, rewrites, bug fixes, zero re-prompting
The model caught its own GLSL compile error from the browser console, fixed it, then programmatically verified its work (leg IK attachment, house dimensions).
Two consumer GPUs, one model, one session, one HTML file. Local agentic coding is right there. 🪶
On 2 consumer cards:
Qwen3.8-Flash-Next 3.05bpw on 2× RTX 3090 + Ryzen 9 7950X + 64GB RAM, running TabbyAPI/ExLlamaV3 with CPU offload.
Concurrency set to 2:
• Decode: 33–75.5 tok/s per request
• PP: 290–776 tok/s on new tokens
• 20K–32K prompts, 83–99% cached CTX Length: 200K
• MTP enabled
#localAI #rtx3090 #ai
🤖 Made with AI
Erman Eroglu retweeted
Crossed 7500 followers!
I'm incredibly grateful for everyone who made this happen.
Expect more posts in local AI, good, bad and everything in between.
Hopefully by next milestone I'll have landed a job 😭
Ternary-Bonsai-2-27B-MLX-4bit
Context : 32K
on my Macbook M2 Max with 32 GB of ram:
Prefill 62 t/s
Generation: around 10 t/s :)
3D Cycling Pelican :)
Dual 3090 + 64GB DDR5 (6000). Local. Qwen3.8-Flash-Next EXL3 2.05bpw (turboderp/Qwen3.8-Flash-Next-exl3, rev 2.05bpw_h4_ng4, 59GB).
#Qwen3_8 #Qwen38FlashNext #EXL3 #ExLlamaV3 #TabbyAPI #RTX3090 #LocalLLM #DeepSeekHarness #MTP
Create a stylish, interactive 3D scene of a pelican riding a bicycle, and display it in the browser.
The pelican should wear a red-and-white cycling cap and sunglasses. Give the bicycle a mint-green vintage frame, and add animated speed lines to emphasize motion.
Let me rotate the scene, zoom in, and adjust the cycling speed. Pay close attention to bicycle geometry, character proportions, and natural pedaling motion. Keep the animation smooth as the speed changes.
Make the page polished and ready for a public demo, with thoughtful lighting, a cohesive color palette, and clean controls.
Test it in the browser yourself and fix any visual or interaction bugs before finishing.
Dual 3090 + 64GB DDR5 (6000). Local.
Qwen3.8-Flash-Next EXL3 2.05bpw (turboderp/Qwen3.8-Flash-Next-exl3, rev 2.05bpw_h4_ng4, 59GB).
Stack: ExLlamaV3 1.4.7 + TabbyAPI. N-gram table in system RAM. MTP on. FP16 KV. 196,608 context reserved. QSA so no quant cache.
Measured, not peak-on-a-slide:
• Short chat (no tools): ~85 tok/s decode, TTFT 2.1s
• Agent @ ~90k prompt, 94% cache hit: ~72 tok/s weighted decode
• Short tool turns: 91–98 tok/s
• Long Write turns: 63–70 tok/s
• Cached prefill: 700–1,220 tok/s
• Cold-ish prefill: ~144 tok/s
• TTFT when hot: 0.4–1.1s
• MTP accept: ~40–90% (long dumps sit at the low end)
• Load: ~15s, 58.5GB weights
2.05 fits both 24GB cards if GPU0 keeps a few GB for prefill. Allocating 196k FP16 + MTP and then sending a 14k harness prompt with max_tokens 32768 will OOM. Cap output, one stream.
Same box previously ran the 27B. This is the 125B Flash-Next path, 6B active, n-gram off the GPU.
#Qwen3_8 #Qwen38FlashNext #EXL3 #ExLlamaV3 #TabbyAPI #RTX3090 #LocalLLM #DeepSeekHarness #MTP
🤖 Made with AI
Dual 3090 + 64GB RAM (DDR5 @ 6000 MHZ) CPU offloading,
Qwen3.8-Flash-Next-UD-IQ4_XS + DeepSeek harness, local.
1 turn · 48 steps · 34 min.
Stopped at 65,535 output (truncated=1) — harness cap, not the model. Context was 65.5K / 262K (25%).
Usage looks like 1.99M tokens. 99% cache hit.
Uncached 23K · cached 1.93M · output 39K.
Last slot @ ~65K: prefill 90 t/s, decode 20 t/s, TTFT 5.7s, draft accept 64%.
Short context was ~25–27 t/s; at 65K it’s ~18–21.
Raising the cap is fine — 262K still has headroom. You’ll pay with slower decode and a fatter KV, not an instant wall.
🤖 Made with AI
Erman Eroglu retweeted
Oggi in Consiglio dei Ministri abbiamo approvato l’abolizione del bollo per le auto di piccola e media potenza, fino a 80 kW, e per i motoveicoli, con un’esenzione per un solo veicolo a cittadino.
Abbiamo scelto di trasformare una parte delle risorse impiegate finora contro il caro carburanti in una misura semplice e strutturale, pensata soprattutto per chi utilizza ogni giorno auto e moto per lavorare, accompagnare i figli, muoversi.
Una scelta concreta: cancelliamo una delle tasse più odiate dagli italiani e continuiamo sulla strada della riduzione del carico fiscale.
Mentre altri parlano di patrimoniali, noi togliamo una tassa sulla proprietà.
How much weights we can offload on a single one of these? Which model now ? 🤣
FYI @ItsmeAjayKV
Question is whether it is for consumers or for data centers? I don’t think many of people can afford a 12 TB one
Micron crams 512 GB memory into a single DDR5 stick, pushing next-gen Intel and AMD servers to 12 TB at 9200 MT/s. 🔗 wccf.tech/1l8oc
Erman Eroglu retweeted
Two RTX 3090s. Qwen3.8-27B. 220 tok/s decoding code at 262K context.
fp8 KV on FlashAttention — vLLM ships it for datacenter GPUs only. @AntonProkopyev built the kernel; the club reviewed, ported and measured.
One tier: qwen3.8 27b 64K → 262K, and faster.
github.com/noonghunna/club-3…
P.S. Club 3090 recipes serve all nvidia gpus. Come say hi on discord.