engineer not a scientist. CEO @zenlytic

New York
Joined April 2010
it’s kind of insane LLMs can’t do model.grow(neurons=) yet
Look closer. This is what “change” looks like inside your brain: 1. Stop rehearsing disaster.
133
2020: Parameter scaling 2022: Data scaling 2024: Test-time scaling 2026: RL scaling 2028: Experience scaling (scale with agent-hours of interaction) 2030: Research scaling (scale with AI effort improving AI)
2
89
remember P(doom)?
69
but they all have quite different architectures
Every. single. enterprise company I talk to is building (or already has) one of these right now
1
243
Agents are the first software ever that gets better instead of worse when you add features. and I think people underestimate how important that is
1
110
why is it that as LLMs get smarter, they write with more LLM phrases? Fable throws out "it's not X, it's Y" like every second sentence
2
1
225
This plot should really be on a time axis Because it's terrifying. This line is NOT rounding off as it approaches 'total network takeover'
Replying to @AISecurityInst
On our cyber range "The Last Ones", GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5, from ~7 months before it.
1
127
It's paperclips. They've validated the paperclip thing.
1
2
110
labs are pushing users so hard from gen 1 (chat) to gen 2 (local agent) why they aren't they pushing directly to gen 3 (cloud agent)?
Did... Codex just overtake Claude Code? 24.5 hours ago Tibo announced 6M active users. this means Codex usage jumped 1M in ~ONE DAY. the last user number we heard from Claude Code was 2M in Feb: latent.space/p/ainews-codex-… more analysis within, but this is very big if true.
1
433
many people are burning sol tokens because they're using subagents incorrectly most tasks actually fall in the top-left, not the top-right
Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we're sharing that it produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour. We're sharing the prompt and proof below. We're excited to see what you all do with Ultra!
1
3
1,458
There are only 3 models: - Luna/High - Terra/Ultra - Sol/High
For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper). - Forget everything below Sol High, use Luna with higher effort settings here - Forget Sol Extra High, use Terra Ultra here - The extra cost of Sol Ultra is probably not worth it over Max
57
105
21
1,838
267,088
so am I the only one who hasn’t been trying GPT-5.6 for a couple months or
4
1,726
Product competition usually end with analog comparisons (eg. phone manufacturers now compete on camera quality) I've wondered what that will be for agents. And it seems to be trending the computer use feature.
187
“the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound.”
2
197
WE NEED THE OFFICIAL PELICAN @simonw
64
many people are missing the distinction between /loops and /goals
60
Even though it's the most important product in the world right now, general purpose agent harnesses are still up for grabs. I've tried all of them and none of them give me everything I need. So far it's basically been an either/or: - The frontier models are finally transitioning from the terminal to apps, but they're still architected for smaller, local tasks. - The harnesses have a great Telegram-style experience, but harder to use on a laptop - Even though we know the really valuable stuff for an agent now (self-learning, compounding knowledge, loops), you have to work really hard to get them out of the box. None of them feel really ubiquitous, and they all feel like they're just scratching the surface of what a smart model can do. I think the one to beat right now is Codex with an always-on server. But that's not going to fly for the non-nerd public. And the product that satisfies the full shape has to live outside Claude/Codex by definition. Until the model wars settle down, we need interoperability and the one thing Codex won't do is work with a different model. This is going to get figured out in the next 6 months. Big prize and still anyone's game.
75
2024: the model is the product 2025: the agent harness is the product 2026: the memory is the product 2027: ???
2
53