@nilensoi
iAccount based inCanada
About this account
- Account based in
- Canada
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Employee-owned programmer cooperative in Bangalore.
Joined May 2013
- Tweets1.9K
- Following474
- Followers1.7K
- Likes348
Pinned Tweet
In 2013, a group of makers got together to find new ways to work together.
A lot has happened since. We recently celebrated our 10th birthday :)
Over the years, we've had the privilege of working with some exceptional organizations and doing work we're proud of. 1/2
nilenso retweeted
Nerding out with @GeoffreyHuntley, @headinthebox and the BAML folks was so much fun! Legends ❤️
Solid birthday treat!
nilenso retweeted
Bangalore machas represent at AIE WF! @AtharvaRaykar @threepointone
nilenso retweeted
Our poster is up in the Expo at @aiDotEngineer WF, come check it out!
github.com/nilenso/context-v…
nilenso retweeted
Oh look who showed up to our booth!
Our poster is up in the Expo at @aiDotEngineer WF, come check it out!
github.com/nilenso/context-v…
nilenso retweeted
Swyx wearing the @tracesdotcom shirt around AIE WF was unexpected and awesome!
@swyx you’ve been very kind, and quite helpful too! Thank you for having us present, and thank you for AIE!
nilenso retweeted
I'm at @aiDotEngineer WF with @AtharvaRaykar, and we'll be presenting our work on context-viewer in the Expo.
Come over and lets talk about trace intelligence, error analysis, empirical rigor, semantic data extraction and continuous improvement!
nilenso retweeted
what if you could have an agent search over your team's session history?
- what did we get done this week?
- should we skillify things we're all doing?
- who is using agents in the most interesting ways?
now you can! our new agent is available for our most active users
nilenso retweeted
So, if you recognize these patterns in your agents, and they don't fit your task at hand, you could steer them accordingly.
Full write-up: blog.nilenso.com/blog/2026/0…
Analysis code and data: github.com/nilenso/swe-bench….
nilenso retweeted
I wanted to check this on newer models, but no SWE-bench Pro trajectories exist for Opus 4.6 / GPT-5.4.
So I pulled @badlogicgames' issue-fixing trajectories and ran the same analysis. Thanks for putting those out in public, they make this kind of analysis possible.
Opus's first edit sits at 47% in your pi sessions, vs 35% for Sonnet 4.5 on SWE-bench. Harness and model differ too, so I can't isolate the prompt's effect, but the shape shifts in the direction you'd expect from the explicit analyze-dont-edit prompt.
I think we can see the effect of the human-steering through explicit analysis / go-ahead / wrap-up cues in this comparison.
nilenso retweeted
super interesting work! glad my open traces on @huggingface allow this!
my workflow is mostly:
- prompt template with injected gh issue url.and instructions on how to annalyze and present results + concise impl plan to me
- i confirm analysis/plan either by knowing or double checking manually. may steer to adjust plan a few times until model knows what to do
- tell model to implement
- check results, steer if necessary
- if all good (type checking, linting, tests, manual code review, manual tests), another prompt template is used to wrap up, i.e. changelog, docs, commit, push, comment on issue, close issue
Replying to @SrihariSriraman
I wanted to check this on newer models, but no SWE-bench Pro trajectories exist for Opus 4.6 / GPT-5.4.
So I pulled @badlogicgames' issue-fixing trajectories and ran the same analysis. Thanks for putting those out in public, they make this kind of analysis possible.
Opus's first edit sits at 47% in your pi sessions, vs 35% for Sonnet 4.5 on SWE-bench. Harness and model differ too, so I can't isolate the prompt's effect, but the shape shifts in the direction you'd expect from the explicit analyze-dont-edit prompt.
I think we can see the effect of the human-steering through explicit analysis / go-ahead / wrap-up cues in this comparison.
nilenso retweeted
I analyzed 730 SWE-Bench Pro trajectories each for Sonnet 4.5 and GPT-5 and turned them into “trajectory shapes”: when they start editing, when they stop, how much they verify, how many steps they spend understanding vs doing.
They have very different work habits.
ALT Two stacked area charts comparing coding-agent trajectory phases over normalized progress from 0% to 100%. Sonnet starts editing earlier, at 35%, and finishes implementation much sooner, by 62%, then spends a long tail verifying after its last source edit. GPT-5 front-loads a lot more reading before it starts editing at 50%, and does very little verification afterward. Sonnet also has to clean up temporary files, while GPT-5 doesn’t.
nilenso retweeted
This was a great opportunity to bridge our industry work at @nilenso with academic research at CMU, and I am very grateful to the co-authors @heathermiller (CMU), Michael Isaac (CMU), and @AtharvaRaykar (@nilenso).
We will be in the Bay Area for the conference soon. If you are building AI tooling or want to talk shop about context engineering, we would love to connect.
nilenso retweeted
We’ll be presenting context-viewer at ACM @CAISconf!
context-viewer is an observability tool for context engineering. It gives structure to LLM contexts using classification by topics, and allows you to compare runs side-by-side. Useful for things like analyzing agent failures, token spends, and evaluating context compaction.
- github.com/nilenso/context-v…
- caisconf.org/program/2026/de…
nilenso retweeted
My takeaways from scanning the Claude Code code for ~45 min this evening:
1️⃣Harness engineering is hard. There's a lot of hard won knowledge in here and plenty of diagnostics to keep the feedback flowing.
2️⃣Harnesses and prompts smooth out model quirks. @SrihariSriraman and I covered this last month, but good to see it verified here. So many conditionals based on model types and specific contexts to deploy to mitigate model weirdness.
3️⃣So much of this is CLI app boilerplate. Fully expect a tool like @badlogicgames's pi to be the foundation for any CLI agent being built today.
I talk about the last point, the opportunity for shared foundations, in a post today: dbreunig.com/2026/03/26/winc…
nilenso retweeted
Did you know claude code has "model counterweights"? These are patches in the system prompt that exist to balance model biases.
These weren't visible earlier, but the leaked code has @[MODEL LAUNCH] annotations that call them out explicitly.
nilenso retweeted
I did a compaction analysis on a couple of recent claude code sessions, and thought I'd share here too.
You can use the link to explore further if you're interested.
These are good compaction examples. I wish I could do the same analysis with some bad examples.
nilenso.github.io/context-vi…
nilenso retweeted
Somehow I didn't fully appreciate how strongly Claude Code's prompt has to fight against the weights to make parallel tool calls. blog.nilenso.com/blog/2026/0…
nilenso retweeted
Something I've been thinking for a while, but finally got to writing it down.
The core thesis is that building reliable AI applications requires a harness to be able to tinker, experiment and iterate, without which the project gets stuck in the prototyping phase.
blog.nilenso.com/blog/2026/0…
nilenso retweeted
Really excited for this one: @SrihariSriraman and I took a deep dive into coding agent system prompts to understand their structure, similarities, and differences. dbreunig.com/2026/02/10/syst…
nilenso retweeted
Replying to @badlogicgames
I've been studying the effect of system prompts in the model + tools + system-prompts + harness stack.
So, I ran the same SWE-Bench-Pro task with Opus+Claude Code, but with different system prompts. One run used Codex's system prompt, and another run used Claude's system prompt.
The workflows on the runs are different, and mirror these kinds of sentiments. You can see the corresponding differences in the system prompts too.
We maybe mis-attributing some of these behaviours to the model, when they're attributable to the system-prompt.