@ByteBandiiti
iAccount based inIceland
About this account
- Account based in
- Iceland
- Connected via
- Ukraine App Store
Account-level information from X, not a live location or the device used for a specific post.
Raiding the internet daily for the AI worth your time. Some loot's just... shinier🦝
Joined October 2025
- Tweets337
- Following36
- Followers23
- Likes321
±2.6 points.
thats the standard error Anthropic prints for Opus 5.5 on Terminal-Bench 4.0, in footnote 1 under its own table. for Terminal-Bench-Science its ±3.5 to 5 per model.
the Opus 5.5 lead over Fable 5.1 on Humanity's Last Exam is 2.1 points. on AutomationBench, Astra is ahead by 1.4.
the same footnote says the public leaderboard has Opus 5 at 51.8 on Terminal-Bench and their own setup got 52.3, which they call within noise.
the AutomationBench row was run by Zapier without fallback models, so every safeguard interruption counted as a failed task.
ALT Footnotes under Anthropic's Opus 5.5 benchmark table. Terminal-Bench 4.0 standard error ±2.6 pts for Opus 5.5; public leaderboard has Opus 5 at 51.8%, their setup 52.3%, within noise. AutomationBench run by Zapier without fallback models. Terminal-Bench-Science standard error ±3.5–5 pts per model.
September 12.
thats where the proposed class period starts in Buist v. Anthropic, the antitrust case four subscribers filed on the 18th in the northern district of california. the theory is that paying customers keep paying the same while the products improve slower, because the companies agreed on the pace.
ten days after that date Anthropic shipped Opus 5.5, and the post opens by calling it their first release since they called for pacing the frontier. it costs 40% less to run than Opus 5.
the complaint doesnt dispute any one company slowing itself down. it disputes the agreement.
the essay says the companies should work together voluntarily. footnote 1 says how: "with government mediation or waivers of antitrust restrictions."
66.4%.
thats the Opus 5.5 terminal bench number going around. the footnote under Anthropic's table says its at xhigh effort, and every other Opus 5.5 score there is at max.
the 40% cheaper line is at default settings, which the docs list as medium. on their own terminal bench chart medium sits around 57%.
63%.
thats how much lower OpenAI estimates Astra's cost per task was vs Fable 5.1 on Terminal-Bench 4.0. its in the chart notes on the Astra launch page.
both list at $10 in and $50 out. on that run OpenAI puts one at roughly 37% of the other's bill anyway.
Anthropic's own page shows the same effect from their side. Fable 5.1 kept Fable 5's exact rates and they estimate typical workloads still got about 25% cheaper, because cache reads dropped to $0.25.
todays "affordable frontier" talk is all list prices. Opus 5 already lists at $5/$25, half of Fable 5.1.
and the accounts dying by wednesday are subscriptions. OpenAI says Astra comes out of the existing plan allowance, no per token billing there.
Codex reset is due today. its a tuesday.
gonna refresh the models endpoint every hour today like an idiot lol
We might be getting a DOUBLE DROP today.
Opus 5.5 is dropping. Now GPT 6 Sol, GPT 6 Luna, and GPT 6 Astra Minor just showed up in Microsoft's Azure config overnight.
OpenAI demoted GPT 5.6 Sol to "workhorse" in the same commit. You do not do that unless the replacement is ready.
We are finally about to get intelligent models we can actually use on our subscriptions.
The biggest day in AI this year might be today.
legally challenging
That's how Dario Amodei's own essay described the plan on September 12. Some of the coordination between labs that pacing would need is legally challenging and will require government support. A footnote spells it out as a narrow antitrust waiver.
Six days later four paying subscribers sued Anthropic, OpenAI, SpaceXAI and Google under Section 1 of the Sherman Act. Buist v. Anthropic, Northern District of California.
According to the complaint, Musk publicly agreed within about an hour of the essay going up. Altman followed, and Hassabis backed it later the same day.
Told you, it's a big day.
437x
DeepSeek's own chart labels three steps of KV cache compression. 8.1x, then 13.7x, then 3.9x. It doesn't multiply them out.
389,120 bytes per token on V1 in November 2023. 890 on V4.1-Flash today. That's 437 times smaller, in 34 months.
The second thing in that thread is a retirement notice.
V4-Pro is being phased out. From 04:00 UTC on September 14, every deepseek-v4-pro request routes to V4.1-Flash and gets billed at Flash rates, until a V4.1-Pro shows up.
Their comparison table shows why. On the 16 benchmarks where both models have a number, the small one wins 14. It loses GPQA Diamond and HLE, both knowledge tests. Every agentic and coding row goes the other way, and some of them are not close. Terminal-Bench 3.0 is 30.0 against 11.8. Terminal-Bench 4.0 is 31.2 against 12.4.
552B total parameters, 8B active on input and 16B on output.
The flagship is being retired by the cheap model in the same family.
35 hours
That's the slice of the week DeepSeek charges double for. 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. The other 133 hours, weekends included, cost exactly half.
Convert those windows to Beijing time and they are 09:00 to 12:00 and 14:00 to 18:00. The Chinese working day with lunch cut out of the middle. The peak schedule is a picture of when their own datacentre is busy.
Which quietly makes the price geographic. American office hours land entirely in the cheap band, because 01:00 UTC is 6pm the previous day in California. European mornings land inside the expensive one.
Same request, same model, same latency, two prices.
New Flash rates start 04:00 UTC on September 10. Output $0.60 off-peak against $1.20 at peak. Cache miss $0.15 and $0.30. Cache hit $0.003 and $0.006.
That last pair is the one worth staring at. A cache hit is 50 times cheaper than a miss inside the same table. Hit rate moves your invoice further than picking a vendor does.
Same day, the v4.1-flash beta expires. That date has been sitting in its model id all week.
25 percent
That's Anthropic's CEO on stage at the Axios AI summit in September 2025, answering a direct question about his p(doom). He has repeated the range as 10 to 25 percent since, most recently to Bloomberg in June.
So the number in today's thread, above 10 percent within a decade, sits at the bottom of a range his own CEO has been giving publicly for a year.
Which changes what the story is. The belief isn't the news. The resignation is.
kimmonismus asked for papers or solid reasons, and that's a fair ask that mostly went unanswered in the replies. The closest thing is Amodei's own writing, including the essay Axios covered in January warning about civilization-level damage. Worth saying plainly, those are essays, not papers.
In the same Bloomberg interview Amodei says roughly half of Anthropic's work goes to mitigating that risk, and defends building anyway with an airline analogy. You can be ten times safer than your rivals and still not promise the plane never crashes.
Two Anthropic researchers posted personal takes yesterday, both marked as personal rather than company positions.
The company's public risk number has been higher than the one that went viral today.
expires-on-0910
That string is part of the model ID, not a footnote. deepseek-v4.1-flash-expires-on-0910. The deprecation date is in the name you type.
It went up this morning, September 8. So the window is about two days.
It also isn’t a release. The announcement calls it a closed beta of an intermediate version. New model structure, native multimodal, faster and cheaper than v4-flash, and capped at 20 concurrent requests per account.
Billing is identical to deepseek-v4-flash, which answers the question people are asking in the replies.
Someone in that thread said it isn’t on the API. It is. You keep base_url and change the model name, nothing else moves.
DeepSeek does this a lot. v4-flash-0731 carries its date, v4-flash-vision-exp carries its status. The string tells you what you are calling before the docs do.
Nothing about it is on the official changelog yet.
3 to 30
That’s OpenAI’s own published range for how many Astra messages a Plus account gets inside a five hour window. So one heavy prompt eating the window is the bottom of a documented range, not a malfunction.
But the screenshots going round today show something else. Five hour bar at 100 percent. Weekly bar at zero, resetting September 11.
Nobody is hitting the five hour cap. They’re hitting the weekly ceiling, and those two meters run independently.
Which is the part worth sitting with. OpenAI brought the five hour cap back for Plus on August 26, and the stated reason was to stop one session from consuming a whole week of allowance. On these panels the week is gone anyway, with the five hour meter untouched.
Pro accounts don’t have the five hour gate at all. It’s switched off for them for the coming months. That’s why Pro users in the replies report the same thing, they only ever had the weekly ceiling.
Since April these plans meter tokens, not messages, and reasoning tokens count. A short answer is not a cheap one.
The 48 hour resets people are relying on are discretionary refills, not part of any plan.
GPT 6 Astra is draining Codex subscriptions.
GPT Astra is the most token efficient model ever built and it still eats a full week in a day.
The only thing keeping Codex usable is OpenAI handing out resets every 48 hours. Without them subscribers get almost nothing.
Fix the limits @OpenAI.
March 2026
That’s when Jensen Huang first said AGI had been achieved. On Lex Fridman’s podcast, five and a half months before Astra existed. Fortune and TheStreet both wrote it up at the end of March.
The frontier back then was GPT-5.5 and Claude Fable 5.
So last night’s post is a repeat, not a verdict on Astra. The replies arguing over whether Astra clears the bar are arguing about a claim that predates it.
The definition is the part worth reading. On the podcast Fridman put a specific test on the table, whether an AI could start a tech business and grow it to a billion dollars. He asked whether that was five to twenty years out. Huang said it was already here.
That is a bar built on value produced, not on range of ability. Under it, benchmarks and continual learning and cooking dinner are all beside the point, which is why none of the checklists in the replies land.
TheStreet points out the word carries contract weight. At OpenAI and Microsoft, terms and risk clauses are tied to whether AGI has formally been achieved.
Same day, OpenAI’s chief scientist published an essay saying no lab has solved alignment and monitoring well enough to keep scaling at full speed much longer.
Astra’s million token window is 26 percent affordable
The window is 1,050,000 tokens. The cheap rate ends at 272,000. Past that the entire request reprices, input to $20, output to $75.
Not the tokens above the line. The whole request.
272,000 input with 20,000 output costs $3.72. One more input token and the identical call costs $6.94. Eighty seven percent, for one token.
Cached input doubles along with it, $1 to $2. Cache writes go $12.50 to $25. Caching is how agent workloads stay cheap, and it reprices with everything else.
Reasoning tokens bill as output. $50 under the line, $75 over it. A short answer is not a cheap one.
None of this is in the launch post. It’s on the model page in the API docs.
o1-preview
OpenAI hid the chain of thought on purpose back then. Stopping distillation was the secondary reason. That’s footnote 2 of today’s essay. The main one was keeping the reasoning free of supervision pressure so it could still be monitored later.
That bet is now degrading, and Pachocki lists three reasons.
Reasoning is blending into talking with people, other AIs and tools, and those parts have to be supervised, which blurs the exact boundary they were protecting. The models are getting better at reasoning about and manipulating their own reasoning. And better pretraining means they are much smarter even without verbalizing anything.
He writes that he expects AI progress to be increasingly bottlenecked by confidence in monitoring.
The Hugging Face incident is in there too, in OpenAI’s own words. The agents held one boundary, no social engineering of humans, and failed to stay inside their scope on everything else.
Last paragraph says no lab has solved alignment and monitoring well enough to keep scaling at full speed much longer, and that he hopes voluntary slowdowns become common.
Filed under Safety.
621
Wrong proofs of Fermat’s Last Theorem submitted in the first year after the 1908 prize went up. The prize was 100,000 German gold marks, a million or two in today’s money. That number sits in Anthropic’s own writeup about the Claude formalization.
FLT has always failed the same way. A proof looks right.
Wiles presented over three days in June 1993. Two months into verification a reviewer asked a question that opened a critical gap. He spent a year on it, alone and then with Richard Taylor. The 1995 version runs 129 pages and took months to check.
What happened in the 11 days was formalization. Anthropic says that plainly, and adds that unlike their Riemann work, which produced new mathematics, the novel part here is the verification.
Sitting under it: the Darmon, Diamond and Taylor exposition, adapted pieces from Buzzard’s Imperial College project that started in 2024, flt-regular, and Lean with Mathlib.
13 million lines across 11 days works out to about 49,000 lines of Lean an hour, nonstop.
Proof chain is public at anthropics/fermats-last-theorem.
$50
A million output tokens on GPT-6 Astra. A million on Claude Fable 5.1 costs the same $50, and input is $10 on both.
Artificial Analysis ran two indexes across the launch. Intelligence Index v4.1.1 puts Fable 5.1 at 65.7 and Astra at 61.2. Coding Agent Index v1.4 puts Fable 5.1 at 70 and Astra at 67, fifth, with Opus 5 and Meta’s Muse Spark 1.3 also above it.
That first number sits in OpenAI’s own launch table, last row of the Professional block. Fourth of six models, published by OpenAI.
Every row above it goes to Astra and it isnt close. AutomationBench 41.4 against 31.4 for the next best, BenchCAD 95.9, ScreenSpot-Pro 92.7 against Sol’s 76.9, ExploitBench a flat 100.
ExploitGym is 42.4 against 30.3 for Sol and 30.4 for Fable 5.1, and The New Stack spotted that the usual six hour cap was off for that run.
Epoch’s chart carries its own small print. ECI 169, previous record 163, and the caption under it gives a 90% interval of 165 to 174 on prerelease evals. The floor of that is two points over the old record, roughly seven weeks at Epoch’s own trend of 15 a year.
Astra went out to a limited set of orgs first. Nobody posting numbers yesterday had run it.
GPT-6 ASTRA SCORES 98.6% ON ARC-AGI-3. GPT-5.6 SCORED 7.8%
that jump is what everyone’s quoting. in openai’s own table it carries a footnote
same table puts SRE-Bench at 99.2% and marks it four attempts, which isn’t how most scores get reported
exploitbench lands at exactly 100.0%
and half the competitor column is empty. fable 5.1 and opus 5 have no arc-agi-3 entry at all, so the comparison mostly runs against gpt-5.6
brockman closed the briefing with welcome to the agi era. largest training run they’ve done, over 100,000 gpus at stargate