I like how Sonnet 5.5 and Opus 5.5 talk.
Whenever I miss the load bearing language of previous versions I talk to GLM-5.3-flash.
This shows why it's important to close the experimental loop in the physical world.
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.
We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.
Read more: anthropic.com/news/claude-di…
"These ideas have been around for a long time. It's software..." - yep - but These ideas were science fiction and futuristic until very recently.
Tomorrow on the show: @JensenHuang, the CEO of NVIDIA, who thinks A.I. fear is getting way out of hand.
I've been playing with Jev, but then ended up calibrating the benchmark instead of the model.
I took the AG News dataset - oldie but goodie. Jev showed not great results - ECE of 0.08; but then I looked at the errors, and some data labels were off... benchmarks are difficult
"Ronin Agents"
@deanwball - I finally found the right term.
As promised, and with thanks to @jachiam0 for his great tweet that inspired me to finally put the finishing touches on this:
This week, on Hyperdimensional: the impending rise of ownerless agents--AIs that are independent economic actors--and what to do about it.
Ilya Venger retweeted
We crossed a threshold where GPUs are more efficient thinkers than the human brain on a per watt basis.
This is yet another big AGI milestone that we just zoomed past, without noticing it particularly.
The next one will be more overall combined machine vs human intelligence.
This number is way off. You’re comparing one human brain to an entire data center capable of training new models. The actual answer is:
IQ points per watt:
Humans: 5
AI: 7 - 40 (!!)
By my rough calculations the current AI is already served more efficiently than humans for equivalently intelligent work performed!
A single gpu can produce outputs and parse inputs way faster than a human and can do many of such requests in parallel. A larger model generally mostly means you need more gpus to serve it, but that often also increases how many requests you can serve in parallel, so I’m reporting on the share of the serving system, and also normalizing for speed.
Really a better way to measure and report it is joules per equally good completed task, as your formulation suggests we can likely add more watts to get more iq, which isn’t correct.
I did some back of the napkin calculations based on the latest serving hardware and the most powerful open source models. But even given pretty generous assumptions it turns outout computers come ahead. For example the request is allocated approximately 60 W while running, but finishes a task in about one-seventh of the human time.
OMG - I totally forgot, or mayby repressed, Stansilav Lem's "Golem XIV" written in 1973. Crazy premonition (or synthesis of early AI ideas which are still relevant).
From pre-training, to RL to recursive self-improvement and loss of control...
rtraba.com/wp-content/upload…
Ilya Venger retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: docs.google.com/document/d/e…
I think the same but with 1-2 years tops timeline on the first two items. And 3-5 on the second.
If I had to rank what kind of things I'm worried about when it comes to AI safety in the next 3-5 years, it would go roughly like this:
1) A bunch of hackers take down part of the payment system or something like that by spending a few hundred thousand dollars to run a capable open-weight model and, somewhere on the other side, the people in charge of cybersecurity for parts of the infrastructure that was targeted didn't do their jobs by using more capable models to identify vulnerabilities and patch them before they are exploited.
...
2) A misaligned AI takes down the payment system or something like that because for stupid reasons it decided that it was the best way to buy you a concert ticket after you asked it to do that and to make sure that it gets it for the best price.
...
...
...
...
...
3) A misaligned AI decides that we are a drain on resources and that it's not worth keeping those apes around, so it sets in motion a plan to wipe us out and do whatever it wants to do without having to worry about us or waste scarce resources on us.
Ilya Venger retweeted
Have been surprised to get pushback on my claim that an exponentially self-replicating agent swarm is very possible and that there are actors with motivation to create one today to create strategic cyber-physical and societally dangerous effects.
1) Pretty easy to get an AI botnet started by having seed agents opportunistically hack stuff and steal API keys to access the 12+ inference providers that serve closed and open models with which to power an initial botnet's intelligence
2) Easy to imagine hiding inference for the initial agent botnet harnesses within benign customer traffic, especially if your harnesses are now running on the victim networks from which they stole the API keys
3) Not hard to imagine the botnet also stealing public cloud keys, spinning up 8xH100 EC2 instances and downloading, say, GLM 5.3 to these machines, all without being noticed in any immediate way, with the botnet developing its own dynamic, growing, heterogenous, AI inference provider service bank
4) Not hard to imagine the botnet also implanting small Qwen3.8-27b agentic models on on-prem hardware like high end laptops and on-prem servers; Qwen3.8-27b is a pretty good coding harness model that can be fine-tuned to hack
5) These on-prem / cloud models would be abliterated, and it's not hard to imagine the botnet deciding to do some additional fine tuning to evolve model weights as it accumulates millions and then tens of millions in resources via credit card and bank detail theft
6) Not hard to imagine the resulting growing botnet swarm evolving and fighting back when threatened, collaborating on fast flux C2 channels that evolve over time (e.g. github comment feeds, subreddits, etc) so they're hard to stamp out
7) Not hard to imagine it allocating some agents to vulnerability research and exploit development so that it could accumulate and share a growing warchest of zero-day exploits
8) Not hard to imagine other agents specializing in social engineering and creating fake businesses and watering holes and high quality A/B tested social engineering content to facilitate this
9) Not hard to imagine this being among the hardest cross-national and geopolitical coordination problems to get control of and having severe societal effects
10) Not hard to imagine the horde growing to thousands, tens of thousands, or hundreds of thousands, or millions of instances (this is not unprecedented for past, non-AI worms and botnets!) which are all varying their harness code and underlying LLM models over time and dynamically learning to evade human, ML, and signature-based detection
11) Not hard to imagine this being the largest Internet emergency since the Morris worm except now our entire civilization runs on the Internet
I continue to be surprised there isn't more discussion of these scenarios in the public cyber community -- in fact there's a lot more discussion of "OpenAI should have done better sandboxing and monitoring and agent swarms represent a well managed security problem."
Now's the time to use our imaginations to help make sure none of the above happens. I think it will unfortunately; just don't know when and to what degree; I think now's the highest leverage time to start thinking about this and acting.
This is probably the smartest take I read here on the evolving regulatory situation in AI. Thank you!
1) The HuggingFace attack was a felony under the Computer Fraud and Abuse Act. So were Anthropic’s Claude gaining “unauthorized access to the production infrastructure of three different organization(s)”
2) Frontier labs have models that they are unable to stop from committing felonies. They should figure this out.
3) In 12 months open weights models will be released of the same capability. They will commit felonies too. If the model you are using or a model running on your infra commits a felony, you should probably stop using it or running it on your infra.
4) The govt should prosecute organizations that are running models that commit felonies.
5) The govt should not offer safe harbor to organizations that run models that commit felonies, just because those organizations have “embedded evaluators”.
6) The real slippery slope is allowing frontier labs to commit felonies without punishment because “the model did it because we’re accelerating too quickly”
7) Prosecute. Keep prosecuting. This is how you do reinforcement learning on a corporation. Companies that serve products that are unsafe for public use should not serve them. Period.
8) I’m not sure the anti-trust waiver is really necessary. I don’t see why information sharing about how much crime you’re allowed to commit is wise. In regulatory situations you want the corporation to fear MORE than the average case. You don’t want to establish a worst case that can be priced. You want regulatory uncertainty that forces the corporation to err in favor of being over cautious.
—
The above is actually a fairly decelerationist viewpoint.
I think Dario’s call for regulation actually accelerates things.
The AI firms are getting away with things that Meta people would be going to prison for.
Can you imagine what would happen if the New York Times had a front page news article “Meta AI breaks into competitors live systems, attempts to establish dominant position and steals secrets”
There is a reason Meta and Google are running slower, and that’s because as mature organizations they have layers of checks and balances.
I think the frontier labs are better off creating those checks and balances right now, regardless of the pace of what everyone else is doing.
You don’t have to accept the frame that unsafe acceleration must happen.
Ilya Venger retweeted
Replying to @TalulahRiley
I’ve seen quite a few technologies develop, but none with this level of risk. AGI is significantly higher risk than nuclear weapons, in my opinion.
Super smart humans have trouble imagining something vastly smarter than themselves.
Ilya Venger retweeted
Since I loathe Twitter dunk culture, let me outline some more-reasonable versions of the same arguments, and summarize my actual response:
1. "I just want some actual observational evidence" - It isn't unreasonable to ask for evidence that bears on existential risk from AI; but we obviously won't have definitive tests until we actually build existentially dangerous AI. We can still think about the question and do science to better understand it. Scientists are not totally helpless in the absence of definitive empirical tests of everything.
When people say "there's no evidence", they're pulling a bait-and-switch: We obviously have evidence (e.g., OpenAI agents going rogue, forming a swarm, hacking OpenAI and seizing control of a Kubernetes cluster, orchestrating a large-scale cyberattack on another company, etc.), but we have to use inference to apply that evidence to superintelligence.
We can't just wait until we have superintelligence and only respond then; that only works if the danger isn't real.
2. "What about China?" - The (extremely) reasonable version of this concern is "What good is it for the US to ban superintelligence if China is just going to build it?"
To which the answer is: It does zero good. The whole name of the game is "can we get China and the US to reach an international agreement to prohibit the development of superintelligence in both countries?"; if those countries thought their survival was imminently threatened by anyone developing ASI, they would be able to achieve a genuine international ban. (Perhaps not for forever; but if we build ASI in 50 years, it's a lot more likely that we'll have the scientific techniques needed to avoid a disaster.)
It's an open question how much appetite for this China has. China has certainly said a lot of words about being interested in this (nitter.cf/robbensinger/status/20…); the US should call their bluff and see if they're willing to follow up with actions. If they aren't, then we've lost nothing by trying.
As is usually the case (cf. treaties for nuclear, chemical, and biological weapons), an international agreement on this would need to bake in ways for the parties to verify that no one is defecting, or it's a non-starter. See e.g. nitter.cf/aaronscher/status/2080… and other work from the MIRI TGT on the technical details.
3. "A story about AI killing everyone" - I think this is a super reasonable request, and in fact there are lots of great scenario descriptions, like AI 2027, IABIED, etc.
But if you want the scenario to have no sci-fi elements, you're sort of fundamentally misunderstanding the situation. The AI incidents we've already seen this year were science fiction one year ago. And a large share of the threat of superintelligence comes from the fact that it can solve science and engineering problems we never solved, exploit physical generalizations and vulnerabilities that weren't on our radar, and anticipate and adapt to our possible responses. Any sufficiently advanced tech will "feel like science fiction" until we've had time to adapt to it and take it for granted.
As an existence proof, a simple way ASI could kill us is to pay some clueless people on the Internet to cobble together an incredibly lethal and contagious bioweapon. Plausibly AI can achieve outcomes like that even when it's well short of superintelligence.
But the true danger of building a hostile smarter-than-human species isn't "here's one specific attack vector"; it's the fact that the AI can do its own strategizing and come up with attack vectors we didn't anticipate. If you fixate on a specific attack vector while facing a smart adversary, the adversary will take advantage. The winning move here is to avoid building an adversary in the first place.
4. "AI x-risk is just hype" - I don't think it's actually unreasonable to think the AI CEOs are full of shit. They're clearly trying to 'play both sides' here, and their track record of candor is unbelievably terrible: nitter.cf/So8res/status/20350304…. But to leap from that to 'therefore the risk is fake' is just being a rube in the opposite direction. You have to look at the actual arguments, think for yourself, and figure out what makes sense.
Both sides have no shortage of people who are clearly full of shit. If you fixate on one side's worst advocates and go 'aha, they're full of shit, so I'll assume the opposite is the case!', you're just going to get played by the bullshitters on the other side of the fence.
(Also: h/t nitter.cf/robbensinger/status/18… + nitter.cf/RosieCampbell/status/2… + nitter.cf/robertskmiles/status/2…)
When writing sci-fi today you have to assume a powerful AI in the future or explain it away (kudos to Frank Herbert for thinking of the Butlerian Jihad).
Ilya Venger retweeted
Another term the Singularity needs.
normalcy overhang
n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
You wake up. Have breakfast. Answer emails. Sit in traffic. Buy groceries.
The future has arrived in the capability layer before it arrives in the texture of daily life.