@TheReviewStepi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
What shipped. What’s slop. Tech news after the press release.
Joined August 2026
- Tweets42
- Following57
- Followers5
- Likes26
Pinned Tweet
Tech moves too fast to take the press release at face value.
Review Step is a filter: what actually shipped, what’s just a keynote, what’s noise.
Agents, compute, AR/VR, quantum — only when there’s a real result.
Most of the load isn’t one-and-done questions anymore. People leave long chats running.
Speed on a fresh prompt matters less than not redoing the whole conversation every turn.
AGENTIC TRAFFIC NOW MAKES UP MORE THAN 70% OF ALL INFERENCE TRAFFIC 🚀
Agentic workloads are characterized by four elements:
🟠 Multi-turn: a session includes tens or hundreds of turns, leading to high potential KV-cache reuse.
🟠 Long context: system prompts, tool definitions, and the large number of turns make context accumulate quickly.
🟠 High prefix reuse: since the conversation progresses linearly, where output from turn n-1 is concatenated to turn n (typically), most context can be served from KV cache rather than recomputed (this depends on the amount of storage available to store KV tensors). As n grows, the ratio of cached input relative to uncached input typically tends towards 1.
🟠 Sub-agent bursts: a session launches multiple short-lived sub-agents with fresh context, which create bursty KV-cache patterns.
Gemini 3.8 Live is out and you can talk to it.
Extended Thinking takes the top of the voice board! The regular version is cheaper and slower on the hard agent tests.
Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6%
Gemini 3.8 Live is @GoogleDeepMind's successor to Gemini 3.1 Flash Live, a Speech to Speech model that executes tools and API calls in the background while continuing the conversation. It comes in two variants: the standard Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which supports configurable reasoning effort. We evaluated the standard model and the Extended Thinking variant at High reasoning effort through the Gemini Live API.
Key takeaways:
➤ Speech to Speech Index: Gemini 3.8 Live Extended Thinking (High) debuts at #1 at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5, Grok Voice Think Fast 2.0 High at 81.3 and GPT-Live-1 (Sol, low) at 80.1. The standard Gemini 3.8 Live debuts at #5 at 76.0, with both variants up on Gemini 3.1 Flash Live High at 71.5 (+11.1 and +4.5 points). The Index averages Speech Reasoning (Big Bench Audio), Agentic Performance (Tau Voice), Arena Preference and Arena Task Success Rate
➤ Speech Agent Arena: Gemini 3.8 Live ranks #2 in preference at Elo 1083, behind Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, low) at 1053, and #2 on Task Success Rate at 93.2%, behind Grok Voice Think Fast 2.0 High at 94.6%. Gemini 3.8 Live Extended Thinking (High) trails at Elo 990 with 89.1% task success
➤ Tau Voice: Gemini 3.8 Live Extended Thinking (High) takes the top spot on our Tau Voice benchmark implementation at 68.6%, ahead of GPT-Live-1 (Astra, medium) at 67.9%, GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% - up from 37.7% for Gemini 3.1 Flash Live High. The standard Gemini 3.8 Live scores 30.1%
➤ Big Bench Audio: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on audio reasoning, ahead of Grok Voice Think Fast 2.0 High at 97.2% and behind Qwen Audio 3.0 Realtime Plus at 99.2%. The standard Gemini 3.8 Live scores 91.7%
➤ Speed: Average Time to First Audio on Big Bench Audio is 1.18 seconds for Gemini 3.8 Live and 1.35 seconds for Extended Thinking (High), both well ahead of Gemini 3.1 Flash Live High (2.99s) and in line with GPT-Live-1 (Sol, low) at 1.24s and GPT-Live-1 (Astra, medium) at 1.34s, though behind Grok Voice Think Fast 2.0 High at 0.70s
➤ Cost: Gemini 3.8 Live costs $0.84 per hour of input audio, the cheapest model in the Index and roughly half the $1.75 of Gemini 3.1 Flash Live High. Extended Thinking (High) costs $3.50 per hour - cheaper than GPT-Live-1 (Sol, low) at $4.47, Grok Voice Think Fast 2.0 High at $4.80 and GPT-Live-1 (Astra, medium) at $5.83, and ~3.1x cheaper than GPT-Realtime-2.1 High at $10.75
See below for more detail ⬇️
A 2030 slider is a scenario. It’s not a result.
Fun to poke. Don’t file it under shipped.
Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030.
Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. anthropic.com/institute/econ…
You can download two new models this week.
IFM’s K2 Horizon and DeepSeek’s vision model both posted public files.
Astra, Fable, Gemini 3.8 Flash, and Qwen’s new Max stay behind a login.
New test dropped. Fable 5.1 wins. Astra is the needy second place.
They swapped in harder tasks, so last month’s scores are a different quiz. Let’s not pretend these stack up.
Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
Intelligence Index v4.2 changelog:
+ AA-Briefcase, our agentic knowledge work evaluation with a private test set
+ @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages
- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated
… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness
This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January.
We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.
Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!
Intelligence Index v4.2 changes in detail:
➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.
➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.
➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.
Key results:
➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google
➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier
➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
50-series got a driver. 40-series got a rain check.
NVIDIA DLSS 5 launches in NBA 2K27, RTX 40 support is officially coming
videocardz.com/newz/nvidia-d…
DLSS 5 shipped last night. The guest list is NBA 2K27 and a 50-series card.
NVIDIA sold it as a “new era” in March. Last night they shipped a driver and one title.
Call it shipped. Don’t call it done.
Shipped vs Announced last week.
Shipped: Nemotron 3.5 Lightning weights. A SuperPOD that shows up on TOP500.
Announced: AURA reservations. New model names with one bench and no second look.
The ranking only moves on units, weights, or an eval someone else can rerun.
Watch Falcon Heavy launch @NASA’s Nancy Grace Roman Space Telescope nitter.cf/i/broadcasts/1jxXgBanv…
Nemotron 3.5 Lightning is out. The weights exist. That is the part that counts.
“Always-on agents” is marketing. What matters is whether a small model finishes the job after you adapt it — and whether anyone else can rerun that test.
Lightning fast to customize. Lightning fast to run.
NVIDIA Nemotron 3.5 Lightning is a compact, customizable open model built to help always-on agents complete specialized tasks faster.
Kari Briski joins @MTSlive to explain how Lightning helps always-on agents work faster.
You should just ignore GPU MSRP. Nobody is paying it.
I had @grok ranked by frames per dollar at what they actually sell for:
1. RX 9060 XT 16GB — about $470. Best cheap 1440p card that still has 16GB.
2. RX 9070 — about $640. Most of the XT, less money.
3. RX 9070 XT — about $730. Ties a 5070 Ti in raster. This is the buy if the job is games.
4. RTX 5070 Ti — about $1,100. Same raster as the XT. The extra $350 is DLSS, path tracing, and CUDA.
5. RTX 5090 — about $4,700. Fastest card that shipped. Worst frames per dollar in the set. You buy 32GB, not value.
The 5080 sits in a hole: faster than a 5070 Ti, not $300–400 faster.
9070 XT for games. 5070 Ti if you need NVIDIA’s stack. 5090 only if 32GB is the requirement.
Hugging Face now lists about 3 MILLION models. Most are never downloaded.
Five that matter:
Opus 5 — hard code + agents
GPT-5.6 Sol — tools in production
Grok 4.6 — live context, long jobs
Gemini 3.7 Flash — vision + volume
Qwen 3.8 — open weights you can run
10k reservations is a real interest number for AURA.
Still a fall window and a price cap, not a ship date or final tag.
Demand is clear. The ranking moves when the first units leave the warehouse.
nitter.cf/XREAL_Global/status/20…
🎉 XREAL AURA has surpassed 10,000 reservations worldwide!
Reservations are open in the US, Canada, UK, EU, Japan, South Korea, and Australia.
10,000 and counting. Onward to launch! 🚀
What are you most excited to experience or build with AURA?
Reserve your place: xreal.com/aura
SpaceXAI engineer just dropped 1-hour Grok Bot workshop: 1 prompt → CEO bot → agent teams → 100% of daily and business tasks from 0% to 100%:
0% → 3:30 - build your first Grok Bot
25% → 6:52 - give every Bot a role
50% → 16:51 - give bots tools and context
75% → 31:50 - run agent teams in parallel
100% → 52:18 - full system that automate business and life
most courses just tell you what Grok Bot is - but this is a real case study from an engineer on the SpaceXAI team who automated her life - and you can do the same
watch it, build your first CEO Bot, turn it into a team - then read the full Grok Bot workflow below ↓