@_cmd8i
iAccount based inAustralia
About this account
- Account based in
- Australia
- Connected via
- Australia App Store
Account-level information from X, not a live location or the device used for a specific post.
Founding engineer. Building agent harnesses and dev tools
Australia
Joined August 2022
- Tweets603
- Following96
- Followers196
- Likes466
Pinned Tweet
Jev from @typesafeai lets you build self-evolving classifiers
Save this, use it in your prod app: track what falls into "other" and let an LLM turn those cases into new categories
I'm not surprised this is happening, given the state of management in big orgs (especially where engineering isn't a core org), where everything is reduced to performance metrics that no longer work in a world where ai writes code and engineers have become validators
I am done with this shit. It is over. The state of engineering right now is horrible. It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking. I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens.
LLM output should serve to present options with pros and cons so humans can make their own decisions, rather than making decisions for them, that said, this isn't how most people use these tools, many outsource their thinking and decision making to a tool trained to produce the best average output for a given input, packaged in confidently helpful prose
engineering is so back
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
to finally teach us how to distinguish vibe coders from software engineers, the universe gave us @typesafeai
Eugene Oldman retweeted
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
I dont understand the frustration here, formatting should not be the model’s concern but the formatter’s
if your codebase looks like this its simply because you dont use hooks and formatters
Eugene Oldman retweeted
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Eugene Oldman retweeted
We've never seen this before.
The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.
Surprising, because:
1. First time ever that OpenAI is #1 on Vending-Bench
2. The best model is no longer the unethical one.
I'm expecting Claude 5.2 in Sep/Oct 2026, followed by GPT-6.1 in Nov/Dec, with Claude 6 in Q2 2027 and GPT-7 in Jul-Dec 2027. For the OpenAI and Anthropic LLM history below, I counted each version once (Opus and Sonnet share an entry when their version numbers match)
prompt engineering
context engineering
harness engineering
loop engineering
graph engineering
> you're here, next up:
invariant engineering (or first principles engineering)
used Karpathy's autoresearch method on pi with gpt 5.6 sol, fully autonomous, with fable 5 supplying ideas every 30 runs and @orca_build as the orchestrator
across ~160 runs it improved our full unit test run time by 2x or -50%, from 68 sec to 33 sec
crazy how super simple it is to run stuff like this now
1. install pi
2. /login to codex using your subscription
3. install pi-autoresearch (find on gh)
4. then literally use this prompt
/autoresearch optimize unit test runtime, monitor correctness, every 30 epochs use orca-cli to ask fable 5 for more ideas
that's it really, mind blown by all the possibilities this unlocks
Eugene Oldman retweeted
It is -- this is the way we will work with agents in business -- well not with just buzz but others too. My own system agensis.io, one of the first if not the first to go public and be talked about in spaces etc.
Soon a few stealth startups like Type, Buzz, Raft started appearing and had clearly been working on theirs for a while and they'll be more as people jump on this like they did with openclaw etc.
THIS is what openclaw tried to do and there's still a lot of work to do to let "normies" understand this and make it smoother and that's what i'm building, for me and if anyone else wants to try it they can.
CLI agent is open-source, and looking at open-sourcing the rest shortly after I get it to a state I'm happy with.
nitter.cf/_cmd8/status/208086716…
Block's Buzz app is being built with the agents from within the Buzz app itself. No wonder, given that the app unlocks the self-evolving agents paradigm for anyone, just a single prompt away
if you're getting 404 when pairing buzz with your self hosted relay via tailscale use the prompt in the reply, give it to claude on the machine you run the relay on, it needs a second server, that server comes with the relay, but it never starts by itself