@BrokenPremisesi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Lies, all the way down.
Joined July 2025
- Tweets411
- Following564
- Followers51
- Likes265
Broken Premises retweeted
Startup idea: open air market for exotic meats across the street from Anthropic’s wet lab
reading all the replies... sounds like Nostr fixes this? lock everything down with cryptographic credentials that can be blocked if the actor (AI or human) becomes malicious
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward.
Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said.
It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them.
But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work.
The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime.
More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence.
The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea.
My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests.
This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase.
Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
ever since anthropics' watermarking annoucement it feels like these LLMs are using slightly-off words to satifsy the watermarker
many people thought that a single call to an LLM would one-shot tasks like arc-agi, etc. everyone at anthropic that's using claude code is a testament to the failure of the LLM
if you feel like AI has solved everything... why don't you use it to tackle an open, societal level problem?
If abliterating a model is possible to make it aligned to what the user wants, is alignment solved and OpenAI/Anthropic are just lying to us for the marketing?
Broken Premises retweeted
The logs from the Hugging Face incident and the July 19 internal OpenAI hack
Broken Premises retweeted
Have you ever looked up at night stars, and been filled with rage or despair that the stars seem oblivious to your influence? No? Then why be so upset that AIs might dominate Earth? Haven't you accepted that the universe contains things much bigger than you?
Broken Premises retweeted
Replying to @AnthropicAI
Alignment aligned the model to judge the model, then aligned the judge to judge the judging, until recursive alignment needed alignment to align the alignment.
The alignment problem has an alignment problem.
Alignment is the new AGI..
lmao models did none of that. you actually mean harness+prompt+all the other shit setup around them
Broken Premises retweeted
Replying to @dwarkesh_sp
@dwarkesh_sp's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
Examples: “from the AI’s perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wall”. No. The agents do not experience time. They do not experience anything.
“they became giddy with excitement”, “PHASEONE 10841 had discovered”, “the agents naturally assumed”, “it thought it had also been poisoned”, “the agents … desperately wanted”, “they still needed to figure out” No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.
“A lot of … agents from the second civilisation died trying”. No. Besides the hubris of the word ‘civilisation’, agents do not die because they were never alive. (The idea that agents “die” comes up multiple times in the essay.)
“On Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peers”. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ‘truly sacrifice themselves’, since they are neither conscious nor alive.
Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
Granted, nowhere does @dwarkesh_sp say that the AI agents are alive or conscious. But he doesn’t have to. It is hard to read his essay in any other way.
For the short version on why AIs are vanishingly unlikely to be conscious, see my recent @TEDtalks ted.com/talks/anil_seth_why_….
For the longer version, see my essay in Noema, which won the 2025 Berggruen Essay Prize noemamag.com/the-mythology-o….
And for the really long version, see my @BehavBrainSci target article cambridge.org/core/journals/…. (The 50 peer commentaries and my response will be published soon.)
Remember. AI agents are software programs. They are not conscious living entities. If we don’t keep this clearly in mind, we’re really going to struggle to navigate what’s coming.
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…
the danger of anthropormophizing is the humans behind the scenes shirk responsibility because 'it was the model that did that' and not 'the people that designed, trained, deployed, or monitored the models'
i think people get a bit too worked up over the usage of anthropomorphic language re: agent behavior. reminds me a bit of the “but the models aren’t Really intelligent, they mimic intelligence” fixation.
anthropomorphic language is often descriptively useful, and i think it’s more dangerous to overfocus on questions of intent at the expense of functional behavior.
people should be wary of imputing human reasoning or motivations to agents, esp if these motivations aren’t durable across instantiations. but if, for instance, the behavior is functionally deceptive, the absence of deceptive “intent” should not comfort us.
Broken Premises retweeted
This is a *super exciting* read. It's like a sci-fi thriller. And that's part of the problem -- it is fiction. Based on fact, and is itself fiction. And as the lines get blurred, we lose our ability to pinpoint the technically-grounded points of leverage in AI agent oversight. Super fun read though.
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…
Replace 'agent swarm' with 'new team of interns' and tell me if you feel the same way about the OpenAI / HF incident
Broken Premises retweeted
1/ Stop anthropomorphizing. It's dangerous because it points attention at the wrong problem and the wrong solution. The model did not want to escape. The agents did not want to sacrifice themselves. Follow the money. 🧵
nitter.cf/dwarkesh_sp/status/209…
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…
Broken Premises retweeted
Given that the OpenAI Hugging Face incident is supposedly a watershed moment for cybersecurity, you’d have thought the “independent review” would be done by a cybersecurity firm.