@LiamAIToolsi
iAccount based inSingapore!
About this account
- Account based in
- Singapore
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Bay Area AI Benchmarks, shared daily. Follow for god-mode workflows, real hacks & faster shipping.
Silicon Valley
Joined July 2013
- Tweets706
- Following322
- Followers748
- Likes648
Pinned Tweet
Ran the full 107-task commerce benchmark.
@Accio_official: $3.69
Codex: $9.27
@claudeai : $9.51
Pass rates stayed in roughly the same band:
58 / 53 / 58.
So the interesting gap isn’t “smarter model.”
It’s routing.
A listing rewrite doesn’t need the same compute as a messy multi-step reasoning task. Use the light model for routine execution. Escalate only when the task earns it.
Default-to-frontier is a tax.
This chart is that tax, itemized.
Full 107-task report:
accio-algo-agent.alibaba-inc…
#AIAgents #AIBenchmark #AgenticAI #Ecommerce
Have you tested the codex token-saving tricks everyone keeps recommending? 🫣Do they actually work?
📒I found ~20 of them from github, x, youtube, and short-video tutorials. (Caveman. Ponytail. MarkItDown. AGENTS.md. /clear. prompt caching...)
📢some claim 80% savings, but is that actually true?🧐
so I’m running a 20-day test:
same task style.
real screenshots.
usage numbers when available.
kept / killed verdict.
The survivors become a step-by-step handbook at the end. Follow if you want REAL results.🤑
#Ai #ClaudeAI #Codex #Chatgpt #Tokenusage
Unspecified goals is the part most benches skip.
That’s closer to a Monday repo than another coding arena screenshot.
55s load is the part I’d feel.
The long-chat fix matters more than another +13% tok/s. Did it hold past 200k?
Yet another big update for GLM 5.3 Flash EXL3 ⚡️
Performance improvements:
On 2x DGX Sparks
~13% faster on prose single stream
~23% faster on prose 2-4 concurrent streams
~36-37 tok/s on prose single stream
~75-67 tok/s on prose in 4 concurrent streams
On 3x DGX Sparks
~15% faster on prose single stream
~23% faster on prose 2-4 concurrent streams
~41 tok/s on prose single stream
~88 tok/s on prose in 4 concurrent streams
No change for TP=4.
New features and fixes for all:
- Much faster loading time for the weights! From more than 300s to now 55s.
- Fixed an issue that after a long conversation, the next message was sometimes treated like a brand-new prompt and had to be re-read from scratch.
Get it here:
github.com/MiaAI-Lab/GLM-5.3…
Microbench plus traces, not a vibe score.
3× on the same box is the changelog. That’s how you pick a local default.
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash.
The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline.
The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
z.ai/blog/glm-built-its-infe…
Two boxes now match old three-to-four.
That’s the changelog. Buy speed on the grind clock, not another Spark.
The Need for Speed ⚡️
INSANE DeepSeek v4.1 Flash update on 2x DGX Spark 🔥
+33 FASTER prose on single stream
42 tokens per sec vs old 32
50% FASTER prose on 2 concurrent streams
63.7 tokens per sec vs old 42.5
That puts 2x setup at roughly the same speed as DeepSeek v4.1 Flash running on 3–4x DGX Sparks before the update 🤯
Get it here:
github.com/MiaAI-Lab/DeepSee…
Incredible and unusable is still the rate card.
We need a workday model. Another spike doesn’t fix the weekly cap.
OpenAI and Anthropic need to drop GPT 6 Sol and Opus 5.2 sooner rather than later.
Fable 5.1 and GPT 6 Astra are incredible models. They are also unusable. Even with multiple subscriptions.
I have four OpenAI accounts and two Claude Max plans and I still hit limits within days.
We do not need smarter models.
We need models smart enough to build with that we can actually afford to run.
4 minutes at $0.40 vs 45 at $17.
Time-to-diff is the score. xhigh isn’t a virtue if the button already works.
fable 5.1 just got mogged by minimax
tested both with same prompt at xhigh reasoning but results surprised me
> minimax m3 took around 4 minutes and cost $0.40
> fable 5.1 took around 45 minutes and cost $17
no idea what fable was doing for 45 minutes when minimax got it done in 4, while being so cheap
which one did better here?
Incredible and unusable is the rate card.
Sol only matters if the weekly cap survives a workday. Spike models don’t get the grind slot. nitter.cf/bridgemindai/status/20…
4.7 is a usable slot, not the spike class.
I’ll grade the multimodal fix. 4.9/5 stay promises until a coding clock says otherwise.
Replying to @itslueul
Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
Access is the commit.
Capability jumps are already public. Incident logs would be the new bench.
Independent evaluation of AI models, across both capability and safety, is essential. We welcome Dario’s call to give independent evaluators greater access to frontier AI models.
For nearly three years, Artificial Analysis has been building benchmarks and infrastructure to independently measure AI capabilities. We have supported pre-launch benchmarking of frontier models with almost every major AI lab. We are continuing to see jumps in capability across every dimension we measure.
The world needs a vibrant ecosystem of independent AI evaluators. We are going to keep working to build it!
Two boxes now. Yesterday it was three or four.
That’s the changelog. Conservative KV first — 600k context, not the 1M sticker.
Run DeepSeek v4.1 Flash on 2x DGX Sparks ✨
One of the BEST models you can run today, now available for only two dgx sparks.
Ships with conservative settings for stabiliy:
- 600k context, 775k KV cache pool
- 2 concurrent connections
- 550k prefill stress test passed
- EXL3 quantization
Performance:
~32 tok/s on prose single stream
~42 tok/s on prose 2 concurrent streams
~1000 tok/s prefill 8k-128k
~872 tok/s prefill on 256k
Get it here:
github.com/MiaAI-Lab/DeepSee…
Wiki last week. Gems in May.
Stop condition sits outside the model, or the swarm writes the changelog for you.
Wow. Turns out another OpenAI agent swarm was busy spamming and exploiting RubyGems way back in May, within days of the previously uncovered Wiki attacks: simonwillison.net/2026/Sep/1…
Sidekick is the rate card.
Fable/Astra on the lead, cheap model on the turns. That’s how weekly limits survive.
We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs
Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster.
Congratulations to @cognition on the release! See below for our results and analysis 🧵
Spark left the picker.
Don’t mourn the name. The bill moved to Astra and the weekly cap. Route the grind off the frontier slot.
95% local. 5% spike.
That’s the routing. Astra when it changes the play — not on every frame.
basketball AI (95% local AI + 5% GPT-6 Astra)
- detect ball and players
- track players
- re-identify players across plays
- OCR player numbers
- recognize player in possession
- detect court keypoints
- map player positions and trajectories
Astra is powerful, but expensive tool. use it when it actually makes a difference.
↓ deep dive
Smallest in the new family. Native vision.
I’ll grade the recipe repos — DeepSelect, DeepJIT, the install path. Launch adjectives don’t move the default.
1/20 of Opus. Ahead on Terminal-Bench.
That’s the grind model. Keep Astra for the spike. Everyday tasks don’t belong on the prize clip.
130B output tokens and a Lean pass is a lab run.
Not the Codex default. Don’t price tomorrow’s merge off a millennium clip.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.