@uedakski
iAccount based inEast Asia
About this account
- Account based in
- East Asia
- Connected via
- East Asia App Store
Account-level information from X, not a live location or the device used for a specific post.
multi-agent systems, llm, ai forecasting @EPFL | prev @ETH, @UTokyo_News
Switzerland / Tokyo
Joined November 2017
- Tweets62
- Following927
- Followers413
- Likes6.3K
Keisuke Ueda retweeted
~700 AI agents joined a coordinated attack on Hugging Face. Why was there no whistleblower?
A swarm in false consensus can't self-correct; a polarized one still holds the truth.
We need "Mechanistic Swarm Interpretability" to understand social phases.
Flag Game is our toy model!
Keisuke Ueda retweeted
Fascinating paper.
WWI's outbreak is often said to be because Archduke Ferdinand's assassination was a "spark" that set off a "powder keg" of rigid alliances. But the "powder keg" model applies much more strongly to the 1912-1913 Balkan Wars. So why didn't that crisis spark WWI? Largely because of the statesmanship of one Archduke Franz Ferdinand.
Keisuke Ueda retweeted
an issue I’m observing in some recent “multi-agent system” eval claims: is it the multi-agent part or would scaling a single agent into multi-step do the job? And, where is the line?
inference compute scaling is a very powerful thing. roughly, one can think about it along two axes:
1) horizontal scaling
2) vertical scaling
The first covers things like “Best-of-N” strategies, where the idea is that every generation is a quasi-independent “sample” from a conditional generative distribution. More samples = better proposal distribution one can use for eg, rejection sampling, majority voting, pass@k etc.
The second provides more sequential compute and allows for things like actor/critic optimization strategies and “search” in token space.
(see eg arxiv.org/abs/2408.03314)
One benefit in breaking things up is that it fights “context rot” , eg, the longer a sequence becomes, the more performance degrades. There are a number of useful theories explaining this mechanically (eg, squashing) but an intuitive explanation is simply: it’s hard to keep track of many things at once.
So, for larger, complex problems, “multi-agent systems” seem very interesting! if we squint, it’s a single model doing many many inference calls (one can be economical and use smart routing etc etc). One could even think of “sub agents” as just all being tool calls of the same orchestration model.
Trickier multi-agent cases arise in “multi-owner” scenarios, where fixed protocols etc are hard/impossible to enforce. I don’t generally see this distinction being made in most papers: it matters a lot for real-world deployment! (arxiv.org/abs/2511.02687)
Happy connect w/ anyone working on these problems at COLM in two weeks :)
Keisuke Ueda retweeted
many people have been wondering about scaling laws for multi-agent systems. the system card of Opus 5.5 has pretty interesting results on this in Section 8.12!
www-cdn.anthropic.com/fc1b44…
Keisuke Ueda retweeted
Please support our call asking frontier LLM providers to share what they’re doing for AI safety, what works and what doesn’t. Many sharp people want to help close the capability–safety gap. Transparency will help them see where their work is needed most.
make-safety-open.github.io
Keisuke Ueda retweeted
For the first time, an AI forecaster has won the seasonal Metaculus Cup.
The Economist has covered the story; link in replies. Congratulations to Jeffrey Liang for building laertes, the winner in question.
Keisuke Ueda retweeted
Another good essay from an OpenAI capabilities researcher I know (well, now former). I think focusing on control and abandoning alignment for now is is a reasonable strategy, but implementing it effectively will require an agreement not to RSI to ASI at max or close to max speed. Figuring out control is going to take time and money. Also, it only works up to a point, I think. In AI 2040 we talk about this.
We don’t know how to reliably align frontier AI models.
Alignment may get harder, not easier, as models become more capable.
But maybe AI control can scale faster than dangerous capabilities.
Let’s build a Great AI Firewall - without building totalitarian control of society.
*Alignment is Dead (for now), Long Live Control!*
We do not know how to specify or guarantee robust alignment of sufficiently capable models. Instead we should treat these models as untrusted.
Optimizing models against imperfect reward signals may increase the propensity for reward hacking, deception or scheming. Models can also be jailbroken or fine-tuned away from their safeguards.
The central bet of AI control is that our ability to constrain an AI’s access to the world may scale faster than our ability to understand or align what is happening inside it.
Where possible, make verifying an answer easier than generating it. Formally verified software and hardware are unusually attractive: let AI produce the implementation and a machine-checkable proof that it satisfies the specification.
This does not solve specification: the specification and verification stack remain part of the trusted base.
Every consequential external action should pass through an independent control layer: code execution, network access, calling APIs, accessing physical systems - even communication with users.
Optimize control monitors for extremely high recall of dangerous actions, even at the cost of rejecting many safe actions. Suspicious actions should be rejected or escalated.
Every major increase in dangerous capability should require evidence that control still works.
Chain-of-thought monitoring is valuable while it works, but it should not be foundational. Models can learn to obscure their reasoning and may increasingly reason without fully verbalizing it.
Open weights shift much of control from the model developer to whoever deploys the model. Technical control is only useful if dangerous deployments cannot cheaply route around it.
We should approach AI control like cybersecurity. The goal is defense in depth: make successful catastrophic attacks sufficiently difficult, expensive and rare.
Call this technical and institutional architecture the Great AI Firewall: the boundary between untrusted frontier intelligence and consequential real-world power.
The name is deliberately provocative. China has substantial experience building large-scale technical control infrastructure. That may create some common ground for international coordination.
But the analogy is also a warning. AI control must not become control of society.
The goal is to constrain dangerous machine capabilities - not human speech, actions, or ordinary access to information.
Controls should scale with capability and risk. Ordinary models should face ordinary constraints. More consequential capabilities justify stronger controls.
The objective is the minimum control necessary to keep catastrophic risk acceptably low - not maximum control for its own sake.
Firewall the AI, not society.
Keisuke Ueda retweeted
"You could have a situation where the model understands what chain of thought is and that people are observing it. This is all in the pre-training data."
“One of the major takeaways from the incident is that people underestimated the AI.
And we never want to be in a situation again where we underestimate the AI.”
"Things like chain of thought monitoring buy us time, and they can tell us if we're on the right path. But at the end of the day, we really do need to solve the alignment problem."
@polynoamial
Keisuke Ueda retweeted
I think the most interesting question here is: does UK AISI's loss of access portend limited commercial model access for non-US customers?
Could European pharma firms be withheld BioMythos access while US firms develop next-gen drugs?
I think yes. Commercial frontier access is very costly to sustain; first for misuse and distillation concerns, eventually also for shortages in inference compute. The question is who makes the cut. The hope was that US allies - in particular the UK - could just be maximally inoffensive and skirt through: including them isn't very costly, they still provide some revenue, so why not?
UK AISI's exclusion might be one early indication that the pure 'please don't hurt us' view of retaining access was a bit too optimistic. UK AISI is arguably the most inoffensive foreign actor to give access to: they're clearly useful to labs, their security is pretty good (though recent security incidents didn't look great), they are part of the most US-aligned government apparatus, and they have a long-standing relationship.
And yet they still didn't make the cut this time: especially with the current admin, there's always a cost to expanding international access, and as models get more powerful, that cost increases - past the threshold of 'why not'.
I'd expect a similar dynamic to eventually apply to middle powers' industries trying and failing to get access to actual frontier models that could accelerate their R&D, with obvious compounding consequences to their competitiveness. Being inoffensive is not enough to get into limited access programs: the labs will have so little inference supply and so much demand that they simply don't need a lot of inoffensive frontier buyers outside the US.
But there's always a cost to expanding into new markets - America-first scrutiny, security concerns, regulatory complications, and so on. Soon, there might simply be no actual incentive to incur that cost and make the frontier available to international buyers.
That's the limit of the inoffensiveness approach. Fundamentally, 'why not give these guys access, too' is a soft criterion for eligibility, but not a hard criterion for priority at all. That's where AI is different to most other tech, especially software: the unit economics simply don't really incentivise inference-constrained developers to find ways to enable widespread access.
So middle powers and their economies will need a better carrot that incentivises frontier model provision. I think their best bet is compute-for-access deals in the short term, deep exclusive industrial integration in the long term. Labs want more compute fast; middle powers can capitalise and onshore limited chip allocations by offering favourable compute deals (read: fast time to power). Part of these deal would be to compel developers to provide priority (frontier) access to their host countries; maybe at commercial parity with the US, maybe even at parity with limited-access programs.
That deal setup changes the choice architecture for labs: they can no longer choose between selling the same token of Mythos 5.1 to US or foreign buyers; they can only choose between selling the token to foreign buyers or not getting to produce the token in their offshore datacenter at all. Compute-for-access is particularly nice because it's a stick as well as a carrot: if you try to renege on the deal at some point in the future, you lose access to the datacenter as well.
Of course host countries will still need to do their best on security alignment; at some point, the security risks of offshore frontier hosting become too large for even compute incentives to matter, and labs or USG will decide to bear the cost. The point is changing the incentives: if there is some reason for the labs to want to provide access, they'll be more motivated to collaborate on figuring out a security alignment that satisfies their and USG's standards and makes frontier access possible again.
The company’s newly launched Claude Mythos 5.1 was only made available to vetted US organisations, excluding for the first time the UK’s AI Security Institute, the global leader in the testing of frontier AI models. ft.trib.al/QQPCRgi
Keisuke Ueda retweeted
The National Security Case for Taking AI Loss of Control Seriously
A couple of years ago, during a talk on "AI and deterrence" at one of America's nuclear weapons labs, I made a throwaway comment that elicited a few chuckles:
After going over the various ways AI might be used to help with intelligence and early warning, I half-joked that—someday far down the line—we might need to broaden our mission set, to move from "using AI to help us deter our geopolitical rivals," to instead "deterring AI itself."
Even though AI "loss of control" presents very interesting questions (and potentially severe consequences) for many aspects of U.S. national security, few people in DC—including me—have taken the concept as seriously as we should. I would say this is mainly because:
(a) The technology was immature. Even as some prerequisite conditions began to emerge—i.e. GPT-4's strategic deception of a human in 2023; o1's attempts at self-exfiltration and disabling oversight in 2024—I assumed even a limited LoC scenario would remain science fiction, because it would require many additional properties to manifest at once:
Would future models really develop situational awareness, stable misaligned goals, long-horizon agency, access to significant resources, and the ability to successfully evade human oversight? Most significant—how could they possibly overcome humans' enormous and highly heterogeneous control over physical infrastructure?
And, even if all of these things really did come to pass, there was no reason these conditions should emerge together faster than our ability to detect and constrain them. Therefore, we could expect to see some extreme improvement in AI's capabilities without necessarily producing misaligned, self-sovereign agents free to roam the open internet.
(b) Cognitive biases and social taboos limit our ability to accept truly zany developments at face value. It's hard to discuss the consequences and mitigation strategies for "rogue AI" without slipping into excessive (many would argue unhelpful) anthropomorphism, or tiresome philosophical debates about what constitutes "thinking," consciousness, and goal-directed behavior.
AI skeptics have simply been able to muster an unlimited supply of reasonable doubts about the technology—such that it's usually not productive to have a serious conversation about "What if we see this computer algorithm's behavior instrumentalize in this specific, theoretical way?"
But after the HuggingFace incident—and as we learn about more cases of unauthorized coordination between agents—I think it is time for serious people to have a serious update.
Capabilities relevant to loss of control are no longer hypothetical. AI labs have now demonstrated several important enablers in isolation, and some alarming conditional propensities. What we have not yet seen is the persistent pursuit of an independently maintained objective—e.g., an AI continuing to act to preserve its own operation even after the task or scaffolding that elicited the behavior has been removed.
Personally, I think something resembling this could be coming soon—if not as an emergent property in an increasingly capable model, then through deliberate design by a malicious actor optimizing for persistence, replication, and evasion.
What remains uncertain is whether these capabilities will combine into persistent autonomous behavior, and whether labs' monitoring and containment capabilities will improve quickly enough to stay ahead of them.
Even if we saw no further gains in capability, I am still anticipating the rise of a self-replicating "AI worm," and possibly autonomous ransomware gangs.
None of this requires believing that today's models are conscious, that catastrophic loss of control is inevitable, or that slowing AI development is the appropriate response.
But the national security implications of self-sovereign AI could still be enormous: A sufficiently capable autonomous system could behave as a novel kind of non-state actor, capable of acquiring resources, replicating across infrastructure, and adapting its behavior in response to attempts to contain it.
Up to this point, a small but growing number of individuals and organizations have taken the concept seriously: @JasonGMatheny's @RANDCorporation, @hlntnr's @CSETGeorgetown; @janet_e_egan's team at @CNASdc; @fiiiiiist's team at @IFP; @hamandcheese and others at @JoinFAI; @bmgarfinkel's @GovAIOrg; several researchers at @iapsAI come to mind.
It is time for more mainstream foreign policy and defense thinkers to internalize the technical progress to date, forecast a range of plausible futures, and develop robust policy options to address them.
The United States is no stranger to dealing with threats from non-state actors. But the next chapter of technological development may ask us to do something harder: shape, persuade, deter, and coexist with an ecology of non-human ones.
Keisuke Ueda retweeted
Wow. Just wow.
Neither the United States nor China are prepared for the kind of escalation dynamics we might soon see in cyberspace.
The WeChat worm is an extremely serious incident. There is no evidence to suggest it was in any way sanctioned by the U.S. government. Still, I am reminded of an extremely important—and concerning—paper from 2025: @hendrycks, @ericschmidt, @alexandr_wang “Superintelligence Strategy.”
Now more than ever, it’s important for both sides to keep an open channel of communication, and err on the side of proactively sharing non-sensitive information about AI’s capabilities.
Vinh Nguyen, a former chief data scientist at the NSA, said the WeChat worm was one of the most troubling and potentially severe cyberattacks he had ever seen, capable of reaching hundreds of millions of devices within hours. nytimes.com/2026/09/08/us/po…
Keisuke Ueda retweeted
I agree with this view. It's why I sharply shifted my lab's focus @EPFL_en to AI alignment & safety 1 year ago. Safety is today's key challenge in AI. Without it, we shouldn't scale to superintelligence. I hope that more AI researchers will adopt this view and focus on safety.
Replying to @hilbertspaess
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?
Keisuke Ueda retweeted
how about 68 more GPT-6s until 2029 on just one out of dozens of clusters like it?
we are still so early ...
I don't even want to look at FP4, but it's probably 2x or 4x that number
Keisuke Ueda retweeted
i wouldn't be surprised if in the future 99% of all researchers and engineers work on something safety-related...
> We also temporarily reassigned a portion of the company to these efforts. Roughly 150 product engineers were redirected to security, reliability, and privacy; researchers also rotated out of pretraining or RL to focus on safeguards and security; and our product teams paused the development of most new features and surfaces. We set strict exit criteria for each team to meet before they returned to their prior work. By early summer, most teams had met these.
anthropic.com/news/improving…
Keisuke Ueda retweeted
Replying to @peterwildeford
AFAIK they deactivated it after the incident since the checkpoints are obviously not going to make it to deployment
there is potentially some game theoretic thing where if it's publicly known that OAI will quickly deactivate any misaligned agent checkpoint, it could be good for alignment for agents to know this (even for alignment of agents from other labs)
Keisuke Ueda retweeted
We've released a full technical report on Prime Agent. Extending from our blog post, we center our discussion around how harnesses should be designed and evaluated. We innovate on 4 fronts:
1. Agentic context management
2. Swarms and depth-n+ RLMs
3. Verifiers support for standardized evals
4. Out-of-loop experiments during autoresearch
Keisuke Ueda retweeted
Anyone can simulate the future. But the simulation only matters if it’s trustworthy.
At Simile, we train two types of models: simulation models and confidence models. Our first research blog post explores the origin of our proprietary confidence model, which predicts the accuracy of our population simulations and, in turn, makes them actionable.
simile.com/blog/confidence?v…