@broadglowi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
matt levine for ai -- ball knower
Joined January 2023
- Tweets2K
- Following521
- Followers1.1K
- Likes1K
1. git clone hermes
2. "Claude, please set up hermes"
So I know Cognition doesn’t care about Little Ole’ Me, and this is just my story, so take it with a grain of salt, but here’s why I don’t think all these crazy scaling revenue numbers by coding agents are going to stick:
We’ve been happy Devin users for a while. Really like the product, and like the company. Still do! This isn’t shade at Cognition, and I’m not upset with them, just a sample of why I worry about the competitive landscape for these cos.
Anyway, our usage grew to >$100k run rate or so. I think one or two of our users even received an email congratulating them for being Top 200 users of Devin.
As usage grew, we added more capabilities and needed a BAA (layman’s terms: an agreement that says they are HIPAA compliant). This required an Enterprise contract (why a piece of paper requires extra money is a topic for another discussion, but it’s very normal…) so we started talking to their sales team.
Basically, we were going to ~1.5x our spend in order to get the BAA. Honestly, that was fine with us. We like Devin!
Our only request was to pay quarterly (already a big concession from our existing monthly payments, given our relatively small size). They declined, unless we committed to EVEN MORE usage.
Before responding, I decided to see how difficult it would be to set up Hermes to do the equivalent work as Devin, and to see what the costs were. I’d host it in our infrastructure and only use providers we already have a BAA with.
Two days.
That’s it. It’s functionally 100% equivalent. Maybe slightly better in some ways, slightly worse in others. We spend a little bit of time futzing with it, but we did that with Devin, too.
The cost is roughly 25% of Devin.
So, two days for $100k or so per year savings. All because they didn’t want us to pay quarterly.
At some point, other companies are going to catch onto this, too. So congrats on this milestone, but keep finding more ways to provide value and keep customers happy. It’s hard to compete.
most interesting thing about this is that it seems quite high-taste. the visual language is more or less 2026-coded. curious how much of that was prompt vs intrinsic to the model's new capability
Ok, thanks for coming everyone, who's up first on the promo docket?
Hey, so this is Alex, he's a TPM on the App Store Ads team, and his manager says he's very solid. He's up for a promo from L5 to L6. Good peer reviews all around.
Ok, can we see his impact report?
Yup - the highlight is that he increased App Store Ad Revenue from $22.2B to $26.8B - a 21% uplift sir.
No kidding! How?
Color change sir, - he deleted one line of code. He removed ad background coloring so consumers can't tell when they're clicking on an ad. Says he's done it at Yahoo, Bing, Google, and Facebook since 1999, and it never fails. The man's built his whole career on making ads slightly less perceptible. Also, he's been working remotely from a chalet in Val-d'Isère and he's got two kids at Harker, so I think he needs the extra $800k.
Well, this the kind of talent we have to foster -- he's a shoo in!
🫤 Apple looks to be removing the background color in App Store search ads, making them easier to mistake for organic search results. Gotta hit those quarterly numbers! nitter.cf/thomasbcn/status/21029…
broadglow retweeted
I’ve seen a couple of posts about this so wanted to demystify. Today, every Muse user gets a free computer in the cloud. It's a real computer, and we’ve designed the security architecture of the Muse Secure VM carefully so you and your Muse can do almost anything you could with a computer sitting under your desk while keeping you and the system safe from threats like prompt injection. We wrote about this at length in our security blog post – security.muse.ai. Activity in the “runtime cell”, which you share with your Muse is unfettered, but sensitive actions are all overseen by the Sentinel, which runs outside of that cell. Similarly, all sensitive secrets - like the passwords you enter into Muse’s secure credential storage - are also stored outside the runtime cell.
The runtime cell gets its own root filesystem (including a full Ubuntu linux image) separate from the host filesystem where your other more sensitive data lives. Because it is isolated from the sensitive stuff that runs on the same box, this means that we can, and do, offer users full visibility and control over the files in the runtime cell. Just as you can when you install Linux on your home computer, you can poke around and see all the files that make the system work - both debian system files and the binaries and data files that implement the parts of Muse which run in the runtime cell.
This was a very deliberate choice - your Muse Secure VM truly is your own computer in the cloud. You can install software in it, write and compile code, use the browser to surf the web: it is your own Linux box that you can operate as you choose with your Muse. Poking around in this computer doesn't give you any privileged access to Meta infrastructure, or to other people's data
If I may geek out a little here for a second… As a kid I loved to take things apart to see how they worked. As a teenager I got into computers and soon found myself drawn to C:\WINDOWS\SYSTEM and the system registry, later Slackware’s /dev/, /proc/ etc – I could see how the system was laid out and as I explored what DLL files and .so files actually did, I gradually became able to meld the computer to my own will.
We’re really proud to be able to put a real computer in millions of people’s hands with a similar level of transparency. We built a file explorer right into the Library tab of the UI. We want you to be able to see the markdown files Muse writes while it thinks about how to serve you better, and explore the internals of the system if you’d like to.
So, when you ask your Muse to show you its entire filesystem, and receive gigabytes of files you’re seeing the full contents of the runtime cell. It’s yours to explore and enjoy!
If you’re not a geek like me, or simply want to download the data that you personally have created directly with your Muse, we added a feature for that too in Settings > Data controls > Download your agent data.
i was having dinner with someone recently who asked why i spend so much time reading & posting on x. she had a pretty negative perception of the platform mostly shaped by the mainstream narrative around it.
my answer was simple. x is where the future gets beta tested.
you basically get to watch ppl build things, talk through ideas, show off stuff, & argue about where everything is going, way way before it reaches anyone else.
there really isn’t another surface on the internet quite like it.
broadglow retweeted
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted.
Is formal verification the future of coding (or at least, bug finding)?
broadglow retweeted
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted.
Is formal verification the future of coding (or at least, bug finding)?
Scale, MSL, Manus - talent like that lets you bet on every hand:
- Image Gen? You can do it
- Open LLM? You can do that too
- Raybans assistant? Yup
- IG Ranking? You bet
- Browser agents? indeed
And people who think building an excellent personal agent is trivial.... lmao. The devil is so in the details, that you need quite a strong, AI-pilled engineering team to do it properly.
Getting AIs to do simple things like manage their own contexts, call the right tools, etc. across the open internet is a hard fucking problem.
We know this because Adept AI was founded on this premise and then shut down. And despite LLMs exsiting for 3+ years now browser agents only worked at the Opus+ level
alexandr wang just shitposted $290B of META market cap into existence - his last company was just 7% of that and he worked on it for like 8 years
how to make $1B in the next year
- incorporate in antigua
- hire an out of work doctor
- take out massive liability insurance
- prescribe GLP-1s & medical ketamine over Muse
- do the 2 years in a white collar correctional facility
- get out for good behavior
- never look back
Instacart is coming to @Muse! Soon, you can connect Instacart, say “Taco Tuesday,” and turn the idea into a cart from your favorite store. Then check out and get groceries delivered to your door. Same Instacart grocery experience, new way to get started. tr.ee/6xWsKG
if you're building a company in a crowded space, you will need to grind 80% your team 9-9-6 to get to feature parity while the remaining 20% meditate and walk in nature to actually innovate and build the core differentiation