FAANG Engineer | $1M soon | Uiuc Grad| Passionate Ai, tech loves coding. Daily studying , growing , Ai Agents, productivity , gym Oxys | Travel to explore

Seattle, Washington
Joined February 2024
kj.staff.engineer retweeted
please don't nerf Opus 5.5 please don't nerf Opus 5.5 please don't nerf Opus 5.5
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
49
83
6
3,273
112,852
kj.staff.engineer retweeted
Replying to @jackfriks
looking into it! it should last ~30 days so this sounds off
2
46
2,447
kj.staff.engineer retweeted
There was a bug with Luna 6 that caused it to have worse vision capabilities than 5.6. That's been fixed!
This quoted post is unavailable.
30
3
269
11,618
kj.staff.engineer retweeted
Did Anthropic nerf Claude Opus 5.5? The first NerfBench results are live. We retested Opus 5.5 and GPT 6 Astra against their own launch scores. Claude Opus 5.5: 99.2% (-0.8% vs launch) GPT 6 Astra: 102.8% (+2.8% vs launch) Verdict: No nerf detected. Opus 5.5's small dip and GPT 6's small increase is within normal variance. More models and more frequent retests are coming. NerfBench only gets better as we collect more data. Which models should we retest next?
204
136
59
2,926
121,189
kj.staff.engineer retweeted
I wanted to know whether combining Opus 5.5 & Astra is worth it, so I tried different approaches on a real task. Astra coordinating Opus 5.5. delivered the best overall result. Opus 5.5 with Astra reviewing offered a better balance of quality & cost. Solo runs weren't far behind & cost much less. This is my own small experiment. Whether the same holds with many threads working in parallel is still an open question.
29
8
1
149
14,846
kj.staff.engineer retweeted
Code freeze isn’t really a thing anymore before releases and in the future the code might even be generated online per request according to some constraints.
692
144
107
5,252
515,688
kj.staff.engineer retweeted
I am most afraid of us eating the productivity gains of agents by just becoming lazier.
310
43
38
2,056
83,237
kj.staff.engineer retweeted
Anthropic did something insane here... I'm using Opus 5.5 all day, and my usage is barely moving right now, the Claude Max plan is insane value legit probably 10-30x more worth than $200 Codex plan
128
60
26
2,058
72,258
kj.staff.engineer retweeted
Resets all propagated. That will be all. Have a fantastic weekend.
2,069
459
550
15,312
1,997,513
kj.staff.engineer retweeted
Opus 5.5 is the best model launch of the year for me... it couldn't have been better the timing was impeccable, they basically made the GPT-6 releases pointless because this is what we're getting: - better, faster and cheaper than Fable 5.1 - with the same feeling Fable had, the one that made it so good to work with - cheap enough to actually run in prod - limits that aren't that bad anymore so a lot more people can afford it now, which genuinely opens up use cases that were out of reach before they put OpenAI on silent mode for a few days with this one
38
8
1
348
21,890
kj.staff.engineer retweeted
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
1,083
1,229
804
34,783
2,562,473
kj.staff.engineer retweeted
o yes… we’re back in action and we’ll reset usage limits for all paid users across codex and ChatGPT work sorry about the brief disruption! (and yes we have a special spare codex when things are down to help us out)
4,264
817
1,321
17,670
4,685,155
kj.staff.engineer retweeted
lots feedback here, many of you are planning yourself & don't need plan mode others prefer the UX of entering a mode where Claude is just thinking & brainstorming with you my plan is to: - make plan mode into a built-in mod - allow mods to add new modes or override shift+tab
we’re thinking of killing plan mode and using the shift+tab hotkey to adjust effort levels I don’t think the models need plan mode anymore, but if you’re a plan mode diehard would love to get your feedback on why
272
44
24
2,172
331,286
kj.staff.engineer retweeted
One of our most requested updates! The team has been working on it for a while, much nicer experience when you hit your limit mid-task 👏
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
178
76
12
2,965
105,529
kj.staff.engineer retweeted
Been a bit quiet here because internal Slack has been hilarious lately and because we are all locked in on DevDay. Tuesday will be fun.
1,393
190
197
9,352
1,743,996
kj.staff.engineer retweeted
I JUST BOUGHT 2 MORE $200 CLAUDE MAX SUBSCRIPTIONS. 4 Claude Max. 3 Codex Pro. $1,400 a month on AI. I did not cancel a single Codex account. GPT 6 Astra is too good to give up. But Claude Opus 5.5 is the best model I have ever used in my life. I cancelled Claude three times this summer. Opus 5.5 brought me back.
154
21
2
761
62,875
kj.staff.engineer retweeted
If you’re looking for ways to try Opus 5.5 and Claude Tag in Slack…
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
13
4
165
31,389
kj.staff.engineer retweeted
Anthropic has no small models that are worth using right now. OpenAI has no large models that are worth using right now. Google has no models that are worth using right now.
560
488
132
14,446
541,550
kj.staff.engineer retweeted
the right way to use model capabilities is not to ship 10x more features to prod it's to spend more time understanding your users, trying experiments, building prototypes, learning about things you don't understand so that you can ship things that actually work
275
615
180
8,223
385,197
kj.staff.engineer retweeted
if you want to make games, 3d generation is a great capability to help you imagine your game come to life but you should figure out how to make a good, satisfying game loop first (talking to myself here)
15
10
5
472
47,740