True
In just the last 90 days:
1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.”
2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.”
3. Grok 4.6 — they’re frontier. “Still not top 3.”
4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster.
5. Grok 4.8 next month.....
The model machine is just starting up.
Grok will be the workhorse of the upcoming agentic era.
Grok 4.7
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task
Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499).
API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
Important to use Grok 4.7 with our Build harness for the best results
X.ai/build
Grok 4.7 works extremely well with our Build harness
X.ai/Build
Grok 4.7 places @SpaceXAI as third, after Anthropic & OpenAI, for agentic coding.
When factoring in that Grok is significantly faster & lower cost, it’s a great choice for your everyday workhorse.
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol
Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort.
Congratulations to @SpaceXAI and @ElonMusk on the release!
Key takeaways:
➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high).
➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.).
➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively.
Other model details:
➤ Context window of 500k tokens, unchanged from Grok 4.6
➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6
➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
Grok 4.7 is a strong combination of intelligence, speed & low cost
Replying to @SpaceXAI
Grok 4.7 works longer on difficult tasks, checks its work more carefully, and comes with our strongest safeguards to date.
Intelligence is improving exponentially
Don’t mess with 𝕏
Last week, @X sued several people who abused Creator Revenue Sharing by operating a coordinated network of accounts, posting inauthentic content to manipulate engagement, and using multiple bank accounts to hide their scheme.
We do not tolerate fraudulent behavior on X -- and will act forcefully to protect our platform and the earnings of genuine creators.
You can read our lawsuit here: transparency.x.com/assets/le…
A Tesla saves its owner from prison!
Boring Company is working on a simple, precursor Hyperloop tunnel between Austin and San Antonio (>200 mph).
Due to traffic congestion, that journey can currently take up to 2.5 hours. @BoringCompany can reduce that to a consistent <30 mins.