FAANG Engineer | $1M soon | Uiuc Grad| Passionate Ai, tech loves coding. Daily studying , growing , Ai Agents, productivity , gym Oxys | Travel to explore
Seattle, Washington
Joined February 2024
- Tweets10.6K
- Following401
- Followers106
- Likes36.5K
kj.staff.engineer retweeted
Replying to @jackfriks
looking into it! it should last ~30 days so this sounds off
kj.staff.engineer retweeted
There was a bug with Luna 6 that caused it to have worse vision capabilities than 5.6.
That's been fixed!
This quoted post is unavailable.
kj.staff.engineer retweeted
Did Anthropic nerf Claude Opus 5.5?
The first NerfBench results are live.
We retested Opus 5.5 and GPT 6 Astra against their own launch scores.
Claude Opus 5.5: 99.2% (-0.8% vs launch)
GPT 6 Astra: 102.8% (+2.8% vs launch)
Verdict: No nerf detected.
Opus 5.5's small dip and GPT 6's small increase is within normal variance.
More models and more frequent retests are coming.
NerfBench only gets better as we collect more data.
Which models should we retest next?
kj.staff.engineer retweeted
I wanted to know whether combining Opus 5.5 & Astra is worth it, so I tried different approaches on a real task.
Astra coordinating Opus 5.5. delivered the best overall result. Opus 5.5 with Astra reviewing offered a better balance of quality & cost. Solo runs weren't far behind & cost much less.
This is my own small experiment. Whether the same holds with many threads working in parallel is still an open question.
kj.staff.engineer retweeted
Code freeze isn’t really a thing anymore before releases and in the future the code might even be generated online per request according to some constraints.
kj.staff.engineer retweeted
Anthropic did something insane here...
I'm using Opus 5.5 all day, and my usage is barely moving
right now, the Claude Max plan is insane value
legit probably 10-30x more worth than $200 Codex plan
Opus 5.5 is the best model launch of the year for me... it couldn't have been better
the timing was impeccable, they basically made the GPT-6 releases pointless
because this is what we're getting:
- better, faster and cheaper than Fable 5.1
- with the same feeling Fable had, the one that made it so good to work with
- cheap enough to actually run in prod
- limits that aren't that bad anymore
so a lot more people can afford it now, which genuinely opens up use cases that were out of reach before
they put OpenAI on silent mode for a few days with this one
kj.staff.engineer retweeted
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
kj.staff.engineer retweeted
o yes… we’re back in action and we’ll reset usage limits for all paid users across codex and ChatGPT work
sorry about the brief disruption!
(and yes we have a special spare codex when things are down to help us out)
lots feedback here, many of you are planning yourself & don't need plan mode
others prefer the UX of entering a mode where Claude is just thinking & brainstorming with you
my plan is to:
- make plan mode into a built-in mod
- allow mods to add new modes or override shift+tab
kj.staff.engineer retweeted
One of our most requested updates! The team has been working on it for a while, much nicer experience when you hit your limit mid-task 👏
kj.staff.engineer retweeted
Been a bit quiet here because internal Slack has been hilarious lately and because we are all locked in on DevDay. Tuesday will be fun.
kj.staff.engineer retweeted
I JUST BOUGHT 2 MORE $200 CLAUDE MAX SUBSCRIPTIONS.
4 Claude Max. 3 Codex Pro. $1,400 a month on AI.
I did not cancel a single Codex account. GPT 6 Astra is too good to give up.
But Claude Opus 5.5 is the best model I have ever used in my life. I cancelled Claude three times this summer.
Opus 5.5 brought me back.
If you’re looking for ways to try Opus 5.5 and Claude Tag in Slack…
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted.
Is formal verification the future of coding (or at least, bug finding)?
kj.staff.engineer retweeted
Anthropic has no small models that are worth using right now.
OpenAI has no large models that are worth using right now.
Google has no models that are worth using right now.
the right way to use model capabilities is not to ship 10x more features to prod
it's to spend more time understanding your users, trying experiments, building prototypes, learning about things you don't understand so that you can ship things that actually work