Discussing AI, robotics & the latest in tech Building @phaseoteam
United Kingdom
Joined April 2025
- Tweets471
- Following224
- Followers51
- Likes42.1K
Sonnet 5.5 improves on Sonnet 5 across benchmarks, in some cases dramatically.
It’s a faster, lower-cost complement to Claude Opus 5.5, strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.
Sonnet 5.5 is here! A massive upgrade over Sonnet 5 and even close to Opus 5.5 for the same cost. All at the same price as GPT 6 Sol! An amazing release from Anthropic!
Also live now on @phaseoteam!
As I reported over an hour ago, Eleven V4 is here!
Introducing Eleven v4 and Eleven v4 Turbo, our fastest and most emotive voice models yet.
Ranked #1 by Artificial Analysis.
Pepsin retweeted
I stayed up til 2am and spent $1,000 benchmarking Jev Router so you don't have to.
Performance on DeepSWE was roughly the same as GPT-6 Astra on low. It costs slightly more, and it took almost 5x longer to run.
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and @typesafeai
The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost.
Here's how it works 👇🏻
Pepsin retweeted
I’ll be honest: OpenAI disappointed me with this $500 plan.
Not because premium tiers shouldn’t exist, but because of what this normalizes.
$500 today.
$1,000 tomorrow.
$2,000 after that.
At some point, “AGI that benefits all of humanity” starts sounding very different when the best intelligence is increasingly reserved for whoever can afford the highest tier.
AI was supposed to become:
→ cheaper
→ more abundant
→ more accessible
Not a luxury ladder.
If the frontier keeps moving upward while access keeps narrowing, AI stops looking like electricity or the internet.
It starts looking like diamonds.
Artificially scarce, increasingly exclusive, and defined by who can afford access rather than who could benefit from it.
Pepsin retweeted
There's some obvious implications and concerns that a more expensive subscription tier introduces.
1. You pay more for the latest and greatest.
Companies choose their subscription plans carefully.
They gauge the relative usage consumed by their services and price the tiers accordingly based on different levels of consumption.
A higher tier implies that they plan to reveal some product that might demand more usage. Without one, it risks the product's intended experience becoming inconsistent or unfulfilling.
This also suggests that usage of this product will be fairly limited in the lower subscription tiers, which can be "solved" by paying for the higher tier.
2. Lower tiers become less important.
Your pricing ladder has become taller. And your customers have distributed themselves across more tiers.
This leads to the Plus plan having even less users now. So there's less incentive for OpenAI to invest time and money in making the Plus plan a better experience.
Inevitably, this might mean worse limits or more restricted features, like how the 5h session limits were reintroduced to save compute.
People will complain, but it's a quieter voice now because there's less users.
3. A sense of betrayal
"Intelligence too cheap to meter".
"AI for everyone".
OpenAI takes much pride in these values. But they feel contradictory with the introduction of a more expensive plan.
If intelligence is so cheap, why should there be a product that would demand a $500/month subscription?
That's not AI for everyone, that's AI for the rich.
Obviously this is all speculation. And I still think that OpenAI has done impressive work in lowering the costs of their best models over time.
But I got quite annoyed that Astra can't sustainably be used even in the $100 Pro plan.
And I worry that a more expensive subscription tier creates a precedent for more expensive products in the future, at the expense of the lower tiers.
OPENAI 🔥: The upcoming ChatGPT Pro Max plan will cost $500.
So far, this will be one of the most expensive AI subscriptions available on the market.
I hope "it will be worth the wait" 👀
This is a shame, $500 a month. Of all the AI Labs, I least expected OpenAI.
I fully understand that AI is costly, and this will allow them to make more money for more compute, but, doesn't this start making AI exactly what it shouldn't? Only the best accessible by the wealthiest members of society?
And where does it stop? $20 first, $200 next, $500 now, what next? $1k, $2k per month? And how does this affect plans below the top tier? Do they get forgotten about?
This technology has the power to solve everything humanity has ever dreamed of solving, and it seems like sadly, it will become a technology where the very best, benefits the wealthy and causes the gap in society to widen at a rate where it becomes impossible to overturn.
What happens in the next few years is critical for the future of our species. We should not let AI go the wrong way.
Looks like a good upgrade, excited to get these onto @phaseoteam for testing!
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet
these models enable creators, developers, and enterprises to create richer, more expressive audio experiences
try them via the Gemini API and in AI Studio: aistudio.google.com/generate…
50% cheaper Sol and Luna! Wonder if this 50% bump in rate limits?
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
A 5 point lead? Wow! What a release! And cheaper! Wow!
Makes me even more excited for Fable 5.5!
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount
Claude Opus 5.5 brings Anthropic to parity with GPT-6 Astra on evaluations like Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic’s lead in agentic knowledge work.
At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points. Anthropic has cut Opus pricing to $4/$20 per 1M input/output tokens (Opus 5: $5/$25) and cache reads from $0.50 to $0.20.
Key takeaways:
➤ Consistent strong performance, with leading scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam 61.4% (previous best 59.1%, Claude Fable 5.1), SciCode 66.9% (63.1%, Fable 5.1), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA. On Terminal-Bench 4.0 it scores 59.6%, level with the leader GPT-6 Astra (xhigh) and +11 points over Opus 5. It remains slightly behind on CritPt, AA-LCR, and GDP.pdf
➤ Leads in agentic knowledge work: On AA-Briefcase, our private frontier knowledge work evaluation, it reaches an Elo of 1822. This is +143 over Fable 5.1, ahead on both analytical quality and presentation, and is the first time Anthropic has reached presentation quality surpassing GPT-5.6 Sol. This evaluation tests whether models can produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Level with Opus 5 on cost per task despite 1.6x the output tokens: Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max)
➤ Four of five effort levels sit on the Intelligence vs Cost per Task frontier: Opus 5.5 max, xhigh, high, and medium all sit on the Pareto frontier, costing less or outperforming other models scoring 50+ (GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5)
Other model details:
➤ Context window: 1 million token context with image and text input support, unchanged from Opus 5
➤ Pricing: $4/$20 per 1M input/output tokens, down 20% from $5/$25 for Opus 5. Cache writes $5 per 1M tokens for the 5 minute TTL, down from $6.25. Cache reads have been further discounted to $0.20 per 1M tokens, down 60% from Opus 5’s $0.50. This is a 95% discount compared to uncached input pricing, up from 90% on previous Opus models
➤ Effort settings: Five effort settings (low, medium, high, xhigh, and max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled
Opus 5.5 is here! Performing better than Fable 5.1, Anthropic say!
anthropic.com/claude-opus-5-…
Pepsin retweeted
Claude Opus 5.5 is here, and it beats Fable 5.1 and GPT-6 Astra in most benchmarks.
Don't worry, I hype for good reason :3
Claude Opus 5.5 is coming! After rumours about 5.1, 5.2, 5.5 is the next Opus model. Release is confirmed for today!
Looks like GPT 6 Sol is coming today from OpenAI. Exited to see if it launches with the 5.6 Sol Promo Pricing and how it performs!
MiMo-V2.6: The Hard Road to Scaling Up RL
MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems.
The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in.
Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category.
There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time.
To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward.
MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
Pepsin retweeted
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:mimo.xiaomi.com/mimo-v2-6
Pepsin retweeted
Second release of the day! @XiaomiMiMo V2.6 is here!
After showing their public RL run to the world, Xiaomi have publically released V2.6 Flash and Pro.
Showing a substantial increase in performance and priced the same as V2.5, and the best performing Open Weights model in the world!