co-founder @stacktrace_ai | platform to run unsupervised intelligence with confidence. Previously airtable, dbrx, pins & twtr.
California, USA
Joined September 2008
- Tweets6.1K
- Following1K
- Followers2.4K
- Likes7K
Pinned Tweet
📣 Quick life update:
Very excited to have started something w/ @vinod and Prashanth -- Introducing Stacktrace (@stacktrace_ai)!
Some context: At Airtable & Databricks, I worked on helping teams move faster without compromising reliability or security. With unattended agents (especially at scale), the problem compounds: they drift, get stuck or take actions you did not intend.
At scale, teams cannot keep watching the work they thought they had delegated.
I want to give agents more responsibility.
However, to do this effectively -- we collectively need better systems & mechanisms that instill confidence on agents. This is exactly why we are building Stacktrace, a platform that can.
1/ See & explain what agents do
2/ Control what they allowed to do
3/ Intervene when they cross a boundary or drift from the task
@vinod, Prashanth & I have spent much of our careers building and operating infrastructure & security systems at scale. We want to help define "what good looks like" for agentic infra: telemetry that explains what is happening, controls to steer execution and the ability to intervene and manage drift. We believe these foundations are critical to running unsupervised intelligence with confidence, and entrust them with consequential work.
We are working with a few trusted design partners running agents at scale, locally or unattended in the cloud. If you are interested to learn more or try out Stacktrace within your organization, DM me.
-mb
"what did you do this week"
Stacktrace Fleet is out‼️
To celebrate, the Team plan ($199/mo) is free for the first month (no CC required). Try it at stacktrace.ai
AI coding agents log every command, input/output and tool call they make. Stacktrace checks those logs against our rule engine in real time, as they are written.
Stacktrace Fleet gives your team one place to see every agent:
→ Inventory: Every MCP server, plugin, skill and hook in use, across all your agents
→ Findings: Detects leaked credentials, destructive commands, vulnerable packages and agents stalled without saying why, with alerts in Slack in real time
→ Policy: One org-wide policy that decides which components your agents may use, applied everywhere at once
**Most importantly, your data stays with you**. Stacktrace runs alongside your agents. It receives only findings and inventory, never prompts or conversation text.
Happened so many times that we built @stacktrace_ai to solve this
(and other random agent quirks regardless of harness or model)
Micheal Benedict retweeted
It was fun presenting on stage to fellow YC community about @stacktrace_ai at the product showcase this week. Excited about the response we got. Time to build!
A perfect example of how to create value for your customers. TLDR have them advocate for your product doing the right thing
@sqs is absolutely right re: spend — I have seen code review agents spend more in a day where ROI was highly debatable.
I know people who spent $12k in tokens and built nothing.
@Westpac spent $12k in tokens on @AmpCode and built a new intranet for their 35,000 employees (vs. ~$3–5M pre-AI cost).
"Amp is unlocking a huge amount of value in engineering" from itnews.com.au/news/westpac-s…
Last few weeks have been a blur. So excited to be joining @ycombinator this fall.
I feel super fortunate for the immense support from friends, family, some amazing angels, partners & previous co-workers as @vinodkone, Prasanth and I embark on this journey to build out @stacktrace_ai
It is true what they say about the valley -- folks are extremely helpful and they will go out of their way to make things happen. YC is an infectious institution. From first time founders to repeats, the ambition is high and the energy is even higher.
One thing that is very obvious is that it is notoriously hard to keep up w/ the pace in this space. Every other day, there is a new model or some inference technique or an eval result that changes what we thought these systems were capable of (case in point: Jev).
Software has never been this non-deterministic, what worked in one i/p space subtly fails in the next. Old techniques of testing, debugging, knowing "how something works" simply break down. More agents are graduating from individual tasks to persistent virtual employees. At some point, hundreds or thousands of agents will work together and humans cannot simply inspect every trace or grok the why & how behind every action.
Echoing what I have been posting for sometime now -- the primary question to answer will be 'Did the agent do the right thing?' More importantly, what evidence proves it?
I do believe that humans will likely spend far more time defining the evidence & applying judgement to it -- perhaps through evals, safety cases or new (tbd) verification techniques.
That is a big part of what we are excited to work on at Stacktrace. Lots to build.
claude can still generate a script that includes `rm -rf` without vetting/validating (in auto mode)
intent vs drift is still such a large problem
Jev really proves that what we really wanted is a super fast with good precision and recall on basic classification tasks
The stats on this tweet is exactly why people can never get rid of JIRA.
People truly don’t understand the hold Atlassian has on enterprise. Every workflow, process, context on tasks, roadmap, etc lives on these products — Atlassian can be THE company “OS”; however, the challenge here is that their customer base is far too segmented and their top paying users are not the ICP for this future.
I dunno Matt, but do wish him best
Hey, I’m Matt. I’m a Director of Product at Atlassian working to bring agents to the millions of builders who use Jira. Prior to getting into product I was a founder and engineer. Now based in SF and grew up in New Zealand 🥝
In an effort to bring more product voices to X, I’m going to start sharing what I’m building and how I think about product craft. If that’s of interest, let’s connect. My DMs are open and I’ll be answering questions in the thread.
the amount of slop (now livestream story/video gen) to make a quick buck is mind-boggling
I sincerely hope we (as a community) continue to value craft, taste and take actual pride in what we build
lot of matic chatter over the last few weeks, quite curious to see what is all the buzz about cc @mehul
i am starting to use more of codex/sol now than opus -- this was not the case a month ago.
I still use opus/fable but have to keep reiterating "be coherent, concise and state the facts".
cost of building agents is close to zero. The real cost is having it do what you want, consistently and that is not a tooling + infra problem.
There is really no product that makes it easy to build powerful evals against your workloads/flows that can help “direct”agents to do what it needs to (while reinforcing right behavior)