engineer not a scientist. CEO @zenlytic
New York
Joined April 2010
- Tweets3.2K
- Following760
- Followers1.1K
- Likes1.5K
2020: Parameter scaling
2022: Data scaling
2024: Test-time scaling
2026: RL scaling
2028: Experience scaling (scale with agent-hours of interaction)
2030: Research scaling (scale with AI effort improving AI)
Agents are the first software ever that gets better instead of worse when you add features.
and I think people underestimate how important that is
why is it that as LLMs get smarter, they write with more LLM phrases?
Fable throws out "it's not X, it's Y" like every second sentence
This plot should really be on a time axis
Because it's terrifying. This line is NOT rounding off as it approaches 'total network takeover'
Replying to @AISecurityInst
On our cyber range "The Last Ones", GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5, from ~7 months before it.
labs are pushing users so hard from gen 1 (chat) to gen 2 (local agent)
why they aren't they pushing directly to gen 3 (cloud agent)?
Did... Codex just overtake Claude Code?
24.5 hours ago Tibo announced 6M active users.
this means Codex usage jumped 1M in ~ONE DAY.
the last user number we heard from Claude Code was 2M in Feb: latent.space/p/ainews-codex-…
more analysis within, but this is very big if true.
many people are burning sol tokens because they're using subagents incorrectly
most tasks actually fall in the top-left, not the top-right
There are only 3 models:
- Luna/High
- Terra/Ultra
- Sol/High
For agentic coding, one can say:
- Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper).
- Forget everything below Sol High, use Luna with higher effort settings here
- Forget Sol Extra High, use Terra Ultra here
- The extra cost of Sol Ultra is probably not worth it over Max
Product competition usually end with analog comparisons (eg. phone manufacturers now compete on camera quality)
I've wondered what that will be for agents. And it seems to be trending the computer use feature.
Even though it's the most important product in the world right now, general purpose agent harnesses are still up for grabs.
I've tried all of them and none of them give me everything I need.
So far it's basically been an either/or:
- The frontier models are finally transitioning from the terminal to apps, but they're still architected for smaller, local tasks.
- The harnesses have a great Telegram-style experience, but harder to use on a laptop
- Even though we know the really valuable stuff for an agent now (self-learning, compounding knowledge, loops), you have to work really hard to get them out of the box.
None of them feel really ubiquitous, and they all feel like they're just scratching the surface of what a smart model can do.
I think the one to beat right now is Codex with an always-on server. But that's not going to fly for the non-nerd public.
And the product that satisfies the full shape has to live outside Claude/Codex by definition. Until the model wars settle down, we need interoperability and the one thing Codex won't do is work with a different model.
This is going to get figured out in the next 6 months. Big prize and still anyone's game.