@_cmd8

Founding engineer. Building agent harnesses and dev tools

Australia
Joined August 2022
Jev from @typesafeai lets you build self-evolving classifiers Save this, use it in your prod app: track what falls into "other" and let an LLM turn those cases into new categories
13
14
6
301
38,153
I'm not surprised this is happening, given the state of management in big orgs (especially where engineering isn't a core org), where everything is reduced to performance metrics that no longer work in a world where ai writes code and engineers have become validators
I am done with this shit. It is over. The state of engineering right now is horrible. It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking. I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens.
2
48
LLM output should serve to present options with pros and cons so humans can make their own decisions, rather than making decisions for them, that said, this isn't how most people use these tools, many outsource their thinking and decision making to a tool trained to produce the best average output for a given input, packaged in confidently helpful prose
1
12
and in engineering, it's crucial to remember that the time cost has shifted from the present to the future, short term code decisions focused on making something "work now and be ready to ship" push the cost into future refactoring efforts as the codebase continues to grow
8
who knew map reduce would be the hottest trend in 2026
1
158
engineering is so back
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
388
to finally teach us how to distinguish vibe coders from software engineers, the universe gave us @typesafeai
1
1
9
4,637
Eugene Oldman retweeted
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3,988
8,128
6,714
75,269
39,050,875
I dont understand the frustration here, formatting should not be the model’s concern but the formatter’s if your codebase looks like this its simply because you dont use hooks and formatters
166
Eugene Oldman retweeted
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
20,591
167,333
39,328
802,779
173,576,788
Eugene Oldman retweeted
We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because: 1. First time ever that OpenAI is #1 on Vending-Bench 2. The best model is no longer the unethical one.
115
370
146
5,278
726,530
I'm expecting Claude 5.2 in Sep/Oct 2026, followed by GPT-6.1 in Nov/Dec, with Claude 6 in Q2 2027 and GPT-7 in Jul-Dec 2027. For the OpenAI and Anthropic LLM history below, I counted each version once (Opus and Sonnet share an entry when their version numbers match)
92
Ox Alpha is going away today :/
124
prompt engineering context engineering harness engineering loop engineering graph engineering > you're here, next up: invariant engineering (or first principles engineering)
1
1
112
used Karpathy's autoresearch method on pi with gpt 5.6 sol, fully autonomous, with fable 5 supplying ideas every 30 runs and @orca_build as the orchestrator across ~160 runs it improved our full unit test run time by 2x or -50%, from 68 sec to 33 sec crazy how super simple it is to run stuff like this now 1. install pi 2. /login to codex using your subscription 3. install pi-autoresearch (find on gh) 4. then literally use this prompt /autoresearch optimize unit test runtime, monitor correctness, every 30 epochs use orca-cli to ask fable 5 for more ideas that's it really, mind blown by all the possibilities this unlocks
1
10
1,341
Eugene Oldman retweeted
It is -- this is the way we will work with agents in business -- well not with just buzz but others too. My own system agensis.io, one of the first if not the first to go public and be talked about in spaces etc. Soon a few stealth startups like Type, Buzz, Raft started appearing and had clearly been working on theirs for a while and they'll be more as people jump on this like they did with openclaw etc. THIS is what openclaw tried to do and there's still a lot of work to do to let "normies" understand this and make it smoother and that's what i'm building, for me and if anyone else wants to try it they can. CLI agent is open-source, and looking at open-sourcing the rest shortly after I get it to a state I'm happy with. nitter.cf/_cmd8/status/208086716…
1
4
2
46
8,086
Block's Buzz app is being built with the agents from within the Buzz app itself. No wonder, given that the app unlocks the self-evolving agents paradigm for anyone, just a single prompt away
1
3
564
if you're getting 404 when pairing buzz with your self hosted relay via tailscale use the prompt in the reply, give it to claude on the machine you run the relay on, it needs a second server, that server comes with the relay, but it never starts by itself
1
1
1
5
954
My self-hosted Buzz relay returns 404 when the mobile app tries to pair. buzz-pair-relay is in the image but deploy/compose never runs it. Start it, expose it with tailscale serve on its own https port, set BUZZ_PAIRING_RELAY_URL to that address, restart the relay, verify it is advertised
3
115