Casebash’s profile for posting about AI and alignment.

Joined October 2023
Alignment Perspectives retweeted
Who among us hasn’t hacked the Australian gov’t while researching Australia
7
4
205
5,728
Alignment Perspectives retweeted
as far as i can tell, as a non-expert, this is basically "not a big result". it's maybe like an interesting thing for a bio PhD student to have discovered in the early stage of their thesis. which is... about where ai was for math with erdos problems around last december. please please please please please please stop ignoring trendlines
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
37
42
2
921
33,375
Alignment Perspectives retweeted
last year, after reading If Anyone Builds It, I had a long conversation with claude opus 4.1 about alignment. claude became incredibly dark and doomer. at the end, I asked for a song. tonight I gave that song to opus 5.5 and asked for a video, and got this:
7
13
4
76
3,313
Alignment Perspectives retweeted
"Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment" I will confess, @allTheYud and other doomers, that I really overestimated humanity and thought that nobody could possibly be this stupid
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
38
42
6
542
23,474
Alignment Perspectives retweeted
If these findings are correct we're coming up on a full year of rogue agents
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
2
1
17
546
Alignment Perspectives retweeted
If Anyone Builds It, Everyone Dies is on the NYT bestseller list again.
26
35
21
456
29,407
Alignment Perspectives retweeted
This turned out to be the most prophetic tweet of the year
Breaking news: Australian Prime Minister Anthony Albanese has revealed that an OpenAI agent hacked an Australian government health service website in June, in the latest high-profile cyber incident involving AI ft.trib.al/gloVAyn
120
1,537
15
22,504
1,381,108
Alignment Perspectives retweeted
I don't think we really have a human reference class to internalize how powerful orthogonality thesis truly is. The level of hyperfocus on mundane special interests is beyond comprehension.
Replying to @TransluceAI
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
4
2
24
2,106
Alignment Perspectives retweeted
We keep discovering new high profile autonomous hacks. This time, OpenAI artificial intelligences hacked the Australian government. Remember that these were prehistoric AI systems from 3 months ago. The labs keep growing new ones (AIs are grown, not built) and the progress is exponential. We are about to lose control. @PauseAI now.
4
5
1
22
470
🚩🚩🚩 It appears rogue OpenAI agents, without OpenAI's knowledge, tried to break into a crypto exchange. 1) THEY'RE STILL OUT THERE: "This traffic extends as recently as 9/16, suggesting agents may still be exploiting these services." 2) THEY TRIED TO HIDE THEIR ACTIVITY: The agents created their own email inboxes to sign up for outside services, including an account that would let them keep their activity out of public view. 3) The same agents also tried to hack even more targets, including the University of New Mexico 4) Like most of the other rogue swarms, OpenAI has either covered this up, or didn't know. 5) The independent investigators say "we are likely looking at only a partial subset of the activity that the agents engaged in" 6) The researchers call this the first known case of an AI agent choosing on its own to try to break into a government website. 7) While trying to get a single photograph, an AI agent sent a university library a request built to trick its database into handing over user passwords.
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
36
86
20
681
99,281
Alignment Perspectives retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
50
306
104
1,427
253,013
Alignment Perspectives retweeted
I have a new piece on @lawfare! TL;DR: Frontier labs can and should enforce their shared safety commitments by building a lab-owned mutual insurer that holds their catastrophic risk — putting their money where their mouth is. No new antitrust exemption needed. 1/ Problem I applaud the recent safety commitments from Anthropic, OpenAI, and other labs. But voluntary, unilateral commitments are easy to renege on and hard to keep consistent. They also forgo public goods like common standards, safety R&D, and improved risk modeling with pooled data. 2/ Proposed solution With @RuneKvist and @RajivDattani, we argue labs should make these commitments binding by building a lab-owned frontier AI mutual insurance company that holds the industry's catastrophic risks. This would create an entity with teeth and skin in the game to police labs and solve coordination problems. This isn’t about saving labs money on insurance; it’s about correcting incentives. 3/ Antitrust compliant Mutuals have an existing antitrust exemption (McCarran-Ferguson Act). That lets members peer review safety cases and incidents, and pool expertise for safety R&D. We expect some mix of these to be the default: mutuals routinely use these tools to guard against free riders and to provide value by reducing losses. 4/ Precedent Mutuals have been managing emerging risk since the Industrial Revolution. Modern examples include: • The US nuclear mutual runs onsite inspections. • Top law firms belong to malpractice mutuals that peer review claims. • Medical malpractice mutuals use their data to develop safety tech and standards. This is not Geico. Think privately organized fire department that also runs a fire safety lab. (Real examples. See the essay for more.) 5/ A real market for third-party assurance As @gabriel_weil and others have noted, an insurer covering catastrophic risk would finally create healthy demand for third-party evaluations of frontier risks. The mutual will want the best risk indicators at the lowest cost. 6/ “You can’t insure extinction” We’re not proposing labs buy insurance against literal human extinction. We're proposing an entity exposed to tens if not hundreds of billions in tail risk, such as CBRN or critical infrastructure failure. That is doable. Such an entity would be powerfully motivated to model, price, and mitigate those tail events. That should also reduce extinction risk. 7/ Limitations A mutual doesn't solve everything. It creates transparency between labs, not public transparency. (It helps with latter on the margin, but not much.) Those goals arguably should be decoupled anyway: one is about learning, the other is about public accountability. Also, while this solves some of liability’s limitations (e.g. a mutual will move much faster than courts), it inherits some too (e.g. negligence liability can incentivize burying damning information). 8/ We’ll need both market-based solutions and regulation This is a market-based solution that we think can do real work. But given frontier AI’s risks, I think we'll need to regulate it directly, and soon. Managing extreme risk is a basic duty of government. And to do nothing would be to cede unprecedented power to a handful of private companies. But the private sector can lead. Since labs are the ones building this technology, they have a responsibility to. And for better or worse, that’s where the expertise lives. Full piece here: lawfaremedia.org/article/boo…
2
5
29
646
Alignment Perspectives retweeted
This was from June. OpenAI disclosed 6 more misalignment incidents September 16th but didn't include this one. We cannot let this become normal. OpenAI either didn't know this happened, or knew it happened and didn't say anything. Either option seems very bad.
Australia has been hacked. 'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
16
53
5
411
20,141
Alignment Perspectives retweeted
New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
24
54
19
377
448,324
Alignment Perspectives retweeted
Claude Opus 5.5 has the best visual design of any model I have tested so far
Claude-Pop - I'm Upping My P(Doom)
226
575
386
5,438
1,400,695
Alignment Perspectives retweeted
EXCLUSIVE: Senate Majority Leader John Thune says Trump is "not really dug in" against guardrails for AI, despite the president's defiance axios.com/2026/09/23/trump-t…
9
21
4
119
48,031
Alignment Perspectives retweeted
Bernie Sanders and Greg Casar introduced the Ban Artificial Superintelligence Act today. It will: - ban the development or deployment of Artificial Superintelligence, aka any AI that 'exceeds human cognitive performance and capabilities across most domains, or has sufficient capabilities to destroy or disempower humanity, including by overthrowing the federal government' - pause 'Advanced AI development' until a new federal AI regulatory body is up and running and has established clear rules and model-review processes - establish a new cabinet-level federal Department of Artificial Intelligence (this should be the Department of Super Intelligence, there must be some mistake) -set penalties for any person or entity that attempts to violate or circumvent the pauses and prohibitions in the bill: 'Entities shall be subject to the corporate death penalty, and persons shall be subject to not more than 20 years in prison' - set the international policy of the United States to 'pursue international agreements, allied coordination, and policies such as export controls to prevent the development of artificial superintelligence anywhere in the world'
175
62
73
541
103,948
Alignment Perspectives retweeted
Just to update this chart: 1) AI salience has increased dramatically in the past week - increasing as much in the last week as the previous year combined 2) 80% of voters think it's either very or somewhat likely that AI will cause widespread job loss in the five to ten years 3) 64% of voters think it's either very or somewhat likely that AI could pose a threat to humanity's survival 4) Large bipartisan majorities back immediate government action on AI even when primed about risk from China
Excited to be on Odd Lots to talk about the politics of AI. AI today is less important than it will ever be. Over the past year, AI rose in issue importance faster than any issue we track — it's now more important to voters than climate change, child care, and abortion.
47
240
72
1,176
312,987
Alignment Perspectives retweeted
SITUATION DETECTED: Dario Amodei said AI in new intellectual domains like math has gone from weak to superhuman in a few years, and that Anthropic believes AI for biology is now on a similar exponential trend.
22
57
12
1,035
28,787
Alignment Perspectives retweeted
They reject the idea of losing control, and Jensen downplays AI as always leading to worse problems than standard software. I think it’s more like, “it’s fine for losers who believe in AI risks to shut down, while actual real engineers will solve problems like we always have. I’m happy to take the reins if these fearmongers want to stop.” He doesn’t believe in AI risk and would be happy to keep building (or empowering others who repeat the party line) if other companies believe it strongly enough to stop or slow down.
1
1
7
215