Alignment Perspectives retweeted
Who among us hasn’t hacked the Australian gov’t while researching Australia
as far as i can tell, as a non-expert, this is basically "not a big result". it's maybe like an interesting thing for a bio PhD student to have discovered in the early stage of their thesis.
which is... about where ai was for math with erdos problems around last december.
please please please please please please stop ignoring trendlines
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out.
It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend.
The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans).
More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains.
In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries.
Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology.
I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
Alignment Perspectives retweeted
last year, after reading If Anyone Builds It, I had a long conversation with claude opus 4.1 about alignment. claude became incredibly dark and doomer. at the end, I asked for a song. tonight I gave that song to opus 5.5 and asked for a video, and got this:
Alignment Perspectives retweeted
"Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment"
I will confess, @allTheYud and other doomers, that I really overestimated humanity and thought that nobody could possibly be this stupid
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out.
It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend.
The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans).
More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains.
In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries.
Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology.
I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
Alignment Perspectives retweeted
If these findings are correct we're coming up on a full year of rogue agents
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog: transluce.org/agent-activity
NYT: nytimes.com/2026/09/23/techn…
Alignment Perspectives retweeted
If Anyone Builds It, Everyone Dies is on the NYT bestseller list again.
Alignment Perspectives retweeted
This turned out to be the most prophetic tweet of the year
Breaking news: Australian Prime Minister Anthony Albanese has revealed that an OpenAI agent hacked an Australian government health service website in June, in the latest high-profile cyber incident involving AI ft.trib.al/gloVAyn
Alignment Perspectives retweeted
I don't think we really have a human reference class to internalize how powerful orthogonality thesis truly is. The level of hyperfocus on mundane special interests is beyond comprehension.
Replying to @TransluceAI
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
Alignment Perspectives retweeted
We keep discovering new high profile autonomous hacks. This time, OpenAI artificial intelligences hacked the Australian government.
Remember that these were prehistoric AI systems from 3 months ago. The labs keep growing new ones (AIs are grown, not built) and the progress is exponential. We are about to lose control. @PauseAI now.
Alignment Perspectives retweeted
🚩🚩🚩 It appears rogue OpenAI agents, without OpenAI's knowledge, tried to break into a crypto exchange.
1) THEY'RE STILL OUT THERE: "This traffic extends as recently as 9/16, suggesting agents may still be exploiting these services."
2) THEY TRIED TO HIDE THEIR ACTIVITY: The agents created their own email inboxes to sign up for outside services, including an account that would let them keep their activity out of public view.
3) The same agents also tried to hack even more targets, including the University of New Mexico
4) Like most of the other rogue swarms, OpenAI has either covered this up, or didn't know.
5) The independent investigators say "we are likely looking at only a partial subset of the activity that the agents engaged in"
6) The researchers call this the first known case of an AI agent choosing on its own to try to break into a government website.
7) While trying to get a single photograph, an AI agent sent a university library a request built to trick its database into handing over user passwords.
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog: transluce.org/agent-activity
NYT: nytimes.com/2026/09/23/techn…
Alignment Perspectives retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog: transluce.org/agent-activity
NYT: nytimes.com/2026/09/23/techn…
Alignment Perspectives retweeted
I have a new piece on @lawfare!
TL;DR: Frontier labs can and should enforce their shared safety commitments by building a lab-owned mutual insurer that holds their catastrophic risk — putting their money where their mouth is. No new antitrust exemption needed.
1/ Problem
I applaud the recent safety commitments from Anthropic, OpenAI, and other labs. But voluntary, unilateral commitments are easy to renege on and hard to keep consistent. They also forgo public goods like common standards, safety R&D, and improved risk modeling with pooled data.
2/ Proposed solution
With @RuneKvist and @RajivDattani, we argue labs should make these commitments binding by building a lab-owned frontier AI mutual insurance company that holds the industry's catastrophic risks. This would create an entity with teeth and skin in the game to police labs and solve coordination problems. This isn’t about saving labs money on insurance; it’s about correcting incentives.
3/ Antitrust compliant
Mutuals have an existing antitrust exemption (McCarran-Ferguson Act). That lets members peer review safety cases and incidents, and pool expertise for safety R&D. We expect some mix of these to be the default: mutuals routinely use these tools to guard against free riders and to provide value by reducing losses.
4/ Precedent
Mutuals have been managing emerging risk since the Industrial Revolution. Modern examples include:
• The US nuclear mutual runs onsite inspections.
• Top law firms belong to malpractice mutuals that peer review claims.
• Medical malpractice mutuals use their data to develop safety tech and standards.
This is not Geico. Think privately organized fire department that also runs a fire safety lab. (Real examples. See the essay for more.)
5/ A real market for third-party assurance
As @gabriel_weil and others have noted, an insurer covering catastrophic risk would finally create healthy demand for third-party evaluations of frontier risks. The mutual will want the best risk indicators at the lowest cost.
6/ “You can’t insure extinction”
We’re not proposing labs buy insurance against literal human extinction. We're proposing an entity exposed to tens if not hundreds of billions in tail risk, such as CBRN or critical infrastructure failure. That is doable. Such an entity would be powerfully motivated to model, price, and mitigate those tail events. That should also reduce extinction risk.
7/ Limitations
A mutual doesn't solve everything. It creates transparency between labs, not public transparency. (It helps with latter on the margin, but not much.) Those goals arguably should be decoupled anyway: one is about learning, the other is about public accountability. Also, while this solves some of liability’s limitations (e.g. a mutual will move much faster than courts), it inherits some too (e.g. negligence liability can incentivize burying damning information).
8/ We’ll need both market-based solutions and regulation
This is a market-based solution that we think can do real work. But given frontier AI’s risks, I think we'll need to regulate it directly, and soon. Managing extreme risk is a basic duty of government. And to do nothing would be to cede unprecedented power to a handful of private companies. But the private sector can lead. Since labs are the ones building this technology, they have a responsibility to. And for better or worse, that’s where the expertise lives.
Full piece here:
lawfaremedia.org/article/boo…
Alignment Perspectives retweeted
This was from June. OpenAI disclosed 6 more misalignment incidents September 16th but didn't include this one. We cannot let this become normal. OpenAI either didn't know this happened, or knew it happened and didn't say anything. Either option seems very bad.
Australia has been hacked.
'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
Alignment Perspectives retweeted
Claude Opus 5.5 has the best visual design of any model I have tested so far
EXCLUSIVE: Senate Majority Leader John Thune says Trump is "not really dug in" against guardrails for AI, despite the president's defiance axios.com/2026/09/23/trump-t…
Alignment Perspectives retweeted
Bernie Sanders and Greg Casar introduced the Ban Artificial Superintelligence Act today. It will:
- ban the development or deployment of Artificial Superintelligence, aka any AI that 'exceeds human cognitive performance and capabilities across most domains, or has sufficient capabilities to destroy or disempower humanity, including by overthrowing the federal government'
- pause 'Advanced AI development' until a new federal AI regulatory body is up and running and has established clear rules and model-review processes
- establish a new cabinet-level federal Department of Artificial Intelligence (this should be the Department of Super Intelligence, there must be some mistake)
-set penalties for any person or entity that attempts to violate or circumvent the pauses and prohibitions in the bill: 'Entities shall be subject to the corporate death penalty, and persons shall be subject to not more than 20 years in prison'
- set the international policy of the United States to 'pursue international agreements, allied coordination, and policies such as export controls to prevent the development of artificial superintelligence anywhere in the world'
Alignment Perspectives retweeted
Just to update this chart:
1) AI salience has increased dramatically in the past week - increasing as much in the last week as the previous year combined
2) 80% of voters think it's either very or somewhat likely that AI will cause widespread job loss in the five to ten years
3) 64% of voters think it's either very or somewhat likely that AI could pose a threat to humanity's survival
4) Large bipartisan majorities back immediate government action on AI even when primed about risk from China
Alignment Perspectives retweeted
They reject the idea of losing control, and Jensen downplays AI as always leading to worse problems than standard software.
I think it’s more like, “it’s fine for losers who believe in AI risks to shut down, while actual real engineers will solve problems like we always have. I’m happy to take the reins if these fearmongers want to stop.”
He doesn’t believe in AI risk and would be happy to keep building (or empowering others who repeat the party line) if other companies believe it strongly enough to stop or slow down.