@sdandi
iAccount based inUnited States!
About this account
- Account based in
- United States
- Connected via
- United States App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Surya Dantuluri retweeted
I am thrilled to announce that @beaconholdings has acquired @haizelabs, with me joining as VP of AI Research.
We started Haize in 2024 to enable anyone to build reliable and safe AI. Through our red-teaming and safeguards work with the frontier labs; our observability, guardrail, and evaluation platform serving the world’s largest enterprises; and our pro bono SMB work, any customer could build AI they trusted with Haize.
Joining Beacon lets us deliver the same trustworthy AI to those who need it most, and those most overlooked by Silicon Valley: the essential Main Street businesses the real world depends upon.
Transitioning these essential businesses through the AI revolution is one of the most consequential responsible AI problems of our time. We couldn’t be more honored to tackle it with Beacon.
Thank you to our customers, investors, and team for the journey of a lifetime. And thank you to Nilam, Goutham, Mark, and the Beacon team for the trust and opportunity.
It’s time to get to work.
a country of services delivered from one data center (or permanent capital vehicle!)
autonomous organizations are important way to evaluate productive capacity models add to the economy and how mature they are interacting with people
working with riley for 6+! years i couldn’t imagine working on something as exciting this with anyone else
Some news: 5 months ago I joined @coreauto, a small research lab focused on new ways for models to learn. I’m also hiring four interns to do something my younger self would've loved.
It started with a pretty unusual project. We set up a handful of small online businesses, and put an agent in charge. We gave each business some initial funding and a seed idea, and off they went, trying to sell real things to real customers while we watched. It’s about as open-ended and unforgiving a task as you can probably give to an agent, with countless decisions to make and effects that take time to become clear.
Each business got its own bank account, a debit card to pay suppliers, a web browser, an email inbox, a domain name, cloud infrastructure and a schedule. Every few hours, the agent woke up and decided what to do. It might update a landing page, adjust an ad campaign, reply to a customer, and then go back to sleep.
The agents built actual storefronts and apps. They bought ads, attracted visitors, and even made some sales. The agents acted as operators that kept a business moving, but we noticed they never really stepped back to strategize about what might move the needle. If people weren’t buying, why? Was the landing page unappealing? Or maybe the offer itself was the problem?
One amusing example was the very first order a customer placed. Our supplier had run out of stock. But it’s a dropshipping business, and thankfully tens of other suppliers carried the same product. Instead of trying another supplier, the agent almost immediately sent the customer a polite apology and attempted to issue a refund. Gah! So close!
The agents could handle much of the work of getting a business up and running. What they struggled with was deciding what to change when it wasn’t working.
That’s what we want to understand. We also want to see what changes when people take the lead. What does a business operator actually do all day when agents can do so much? What new work does this allow them to do? Which decisions can’t be delegated?
—
So this winter we're hiring 4 interns as Associate General Managers. Associate Product Manager internships are programs that give early career generalists a product to own. We want to try something similar, but with entire businesses.
You’ll create and grow businesses, with agents handling the routine work so you can focus on figuring out what matters. Each AGM gets a budget and our agent infrastructure that’s designed for this. Which ideas to try is up to you. Put more money behind what finds customers and keep refining.
Where you step in tells us what agents can't do yet. You'll be in our SF office with me and people I’m so fortunate to work with, like @_arohan_, @joannejang, @MillionInt, @marksaroufim and @juliacvillagra
We don't care much where you went to school or what your resume looks like. We’re looking for enterprising people with a drive to create, who’ve made money on the internet before and already use AI extensively to enable their ideas. If that sounds like you: coreauto.com/AGM
Surya Dantuluri retweeted
I don’t post much, but this felt worth unfolding :)
For the past year, I’ve helped shape the core software experience for iPhone Duo, especially multitasking.
A lot of it came down to tiny decisions, the kind that, when they’re right, just disappear into how the product feels to use.
Very grateful for the team and everything I learned along the way. So happy to finally see it out in the world!
there are few research blogs worth reading and parsed.com is one of them
Today we're launching Base Labs, a research lab by Baseten. Our mandate is to make open-source as useful as possible, and our one rule is that we publish without exception, including what fails.
Up until pretty recently I thought the way to get the world onto open models was to train them for one company at a time. @mudithj, @maxkirkby and I cofounded @parsedlabs on that bet, @baseten acquired us, and we spent the last year running their training team doing it for customers one by one. Every one of those engagements taught us something new about how models learn, forget, specialise and get cheaper, and almost none of it got spoken about. Sadly, in general that's the field's default in that the people who know the most about training have the least freedom to say it.
To be clear I don't think the closed labs are the villains here. They get to new capabilities first, which buys the rest of us time to harden the world before that stuff is everywhere, and they're the ones paying to find out what's actually possible. But I'm fairly convinced the only real advantage they have is data and scale, and their incentives point squarely at the frontier. You can't do slow, public science on how these things learn when your job is the best model by end of quarter. Someone without that pressure has to, and there are very few of those someones around. Hence Base Labs.
We have the broad remit of making open-source as useful as possible and our one rule is that we publish without exception, including what fails. The first problem is continual learning, which I have come to think is several problems wearing one name. We are also working on the open RL environments and data that open models need and currently can't get, because we have to aggregate data with the same ferocity everyone's been talking about aggregating compute. Plus a bunch of other stuff I'm genuinely excited about, eg a safety stack people can run on top of open deployments, and performance research so these things are cheap for everyone to serve.
I still think open and closed coexist, and that's the good world. It's just that coexistence isn't free, someone has to actually do the work, and this is basically what keeps me up at night. We're hiring researchers, engineers and fellows. Come help distribute the mandate of heaven!
under appreciated how hard ideating and rendering it into reality is
dar is the greatest at this. incredibly humbling
this corpus of ~.5m tweets was taken over the course of 4 months with the earliest tweets coming from a twitter export of all my bookmarks
many should train their own models and raise as much money as they can sdan.io/blog/intelligence-ar… or none at all
New blog post:
Cursor, Devin, and every app effectively are RL environments. Every session is "free" rollout for training. Serious AI cos will start training their own models soon. Not just for margin but Token Factor Productivity - economic value provided to users per $1 spent; which is why CC and Codex on sub plan is so good and retains well. Hard to do for apps not running inferencing at-cost. More:
sdan.io/blog/training-impera…
building personal software has saved me infinity amount of time and has improved my life experience by an order of magnitude i can’t wait until we have a country of services delivered from one permanent capital vehicle
Surya Dantuluri retweeted
In just over 2 weeks - I head back and hope to see a mom and her 3 cubs survive another winter in Alaska 🤞
“study Long Lake”
Solving the science of asset selection in a future (or indeed the present) where every company is a "Context Acquisition Company" is the real frontier.
I love that everyone is getting around to the idea that the secrets (scarce context) currently illegible to/hidden from computers (human or machine) are everything. Now the next leap for people to make is that the science of sourcing, selecting, and monopolizing that context (really THE ASSETS that produce it) is everything. If AI progress is a function of compute and data (most algorithmic progress is really just data progress; h/t @BerenMillidge, @_kevinlu, @mentalgeorge, @GarrettLord, etc.), then every company is going to have a context desk just like they will (or already do) have a compute desk. The difference is, CONTEXT IS NOT FUNGIBLE. Most context (both that exists right now and that will be created in the future) will be completely commodity beta. Winning will be about getting to and instrumenting the right asset (context production factory) first. And yes, there are right and wrong answers. To do this kind of asset selection well requires an extremely scarce meta-capability: the ability to coordinate the right kind of access and the right kind capital at the right time.
These assets (and the secrets within them) are structurally difficult to access, evaluate and instrument. They are not floating around in banked processes, to be frictionlessly purchased on listed exchanges, or willingly coming through Mercor or Handshake's expert portal. (Yes, a context production asset can be (very often is) a single person or collection of people.)
When @WillManidis talks about a Deal Guy Yuga, what he means is that there are people who have deeply internalized the fact that at the limit, in a world of infinite intelligence, access to/monopoly on the right permissioned data streams is all that matters. Getting yourself to a position (meta-access, meta-capital) where you have the ROFR on those permissioned data streams, means being a generational Deal Guy. This is a very different and specific kind of "Deal Guy" though.
Knowing which asset(s) are going to give you the right context to create, compound, and commercialize the best vertical world model now and into the future is the new form of security analysis. But the triple-exceptional combination of domain expertise, meta-access, and technical ability that’s required to execute this new security analysis effectively is scarcer than the talent at quant firms, YC combined, and dare I say, the labs, combined. Palantir understood this and it's why they focused on getting root-access (or something close) to the "highest-status" institutions, and the data streams they produce, first. If you have the talent that can get access to and create value within those institutions, everything else should be a forgone conclusion.
If you want examples of the teams that (I believe) actually understand this new science of asset selection and long term value capture in a world of infinite intelligence, study Long Lake and @formationbio. They know and have known that it's all about being able to get the right asset (context), in the right market, with the right team (machine and human) first. These two companies are very far ahead on the scientific frontier of context acquisition.
GC backed Long Lake last year. Do you think it’s a coincidence that Long Lake chose to work with General Catalyst? My bet is that Long Lake knew they wanted to acquire Amex GBT before they partnered with GC, and that they partnered with GC because Ken Chenault (the ex-CEO of Amex) is General Catalyst’s Chairman. That gave them the right access at the right time to a very valuable context asset (Amex Global Business Travel)
A superhuman vertical-specific Elon operating every company means market leading monopolies in every single slice of the unstructured economy. The thing is you have to build this superhuman Elon while flying the plane. You can't build this superhuman Elon without the very specific context that operating specific assets in the real world gives you. In fact, there's only one stream of context that was able to produce human Elon! Knowing which context stream is likely to do the same a priori is so extremely difficult, but probably possible.
I’ll let you intuit why Amex GBT is both most likely to be the market leading monopoly if it were operated by the superhuman Elon of business travel and why it’s also the most likely to produce the context to build that superhuman Elon.
The labs of course are very large acquirers of context at present and I think they will continue to play and improve their capabilities here. Through their deplyoment companies, they have already chosen the PE funds that they deem to be the best Context Acquisition Funds. Through in-house deployment focus on Life Sciences they have chosen the vertical they see as containing the most valuable context producing assets. They will acquire very seemingly unrelated companies and will acquihire very interesting people just to get tokens, they will create a Context Acquisition Fund of Funds. But it's not a foregone conclusion that they become the best performing context acquisition companies. Or that they even view it this way. And that presents an opportunity for anyone that does.
frontier intelligence requires being out on the frontier jobs.ashbyhq.com/long-lake
the industral revolution made goods abundant by turning labor into capex and similarly ai is doing the same to cognition: everything software-shaped is marginal inference cost
yet historically America is not software-shaped, its not tool-shaped, its a country of services bundled in to limited liability organisms that trade trust for distribution and the right to be accountable for other humans
what the labs have built -- tool-shaped, inscrutable weight files that echo consciousness, locked behind a series of API calls -- provides ever-growing leverage on digital labor. yet the rest of the economy is not inscrutable; it's owned, operated, liable, naturally anthropogenic of how people manage bureaucracy and drive capitalism.
"You can outsource your thinking, but you can't outsource your understanding"
the core understanding of how a firm runs with its situational judgement, awareness for a customer under economic pressure, its duty for that customer, historical context on how to serve customers better -- was not previously trainable (we're now seeing data co's buy slack logs, emails, etc). synthetically generating this data even by paying someone to roleplay as an analyst or physician; making them feel like their service is being converted into training exhaust breaks the core trust an institution has fought so hard to earn.
Understanding itself is what makes a firm an institution.
in this tweet i also want to poke at the "super stack" -- we have superintelligence, superalignment, yet superdiffusion: intelligence saliently diffused across the economy, doing inter-planetary backpropogation on claims resolved, minutes saved, margin expanded, where the policy is rewarded with USD is a core third pillar humanity needs to safely monitor; focusing on accelerating human capability and expanding human agency with superaligned, superintelligent models that are economically-aligned.
imagine backpropagating across a country.
the only place to run the full super stack is inside an AI rollup. AI rollups inherit trust, build the surfaces of work, and route intelligence to seek exponentially harder economically-relevant reward signals -- accelerating human operators, expanding human agency, and enabling recursive diffusion across the firm's economy. AI rollups are the full-super-stack institutions that turns tokens into labor.
recursive self-improvement will bring us a country of geniuses in a datacenter. but for America -- who will birth the country of services, if not an institution?
This quoted post is unavailable.
this took me longer than expected to write, appreciate @yunyu_l @khushkhushkhush for reading and reviewing
if you want to work on superdiffusion -- reach out