Benjamin Hilton retweeted
We @AISecurityInst performed pre-release alignment testing of Astra.
We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
Benjamin Hilton retweeted
The UK’s AI Security Institute is an organisation close to my heart. Back in 2023 I was involved in its conception and creation. It’s gone on to do incredible things. A genuinely world leading org that shatters the myth that government can’t do things.
So I couldn’t be more delighted to say that I’ve been appointed as AISI’s next Director, working alongside the brilliant @NateBurnikell as its new Chief Strategy Officer.
AISI is the world’s most respected organisation for understanding the risks from powerful AI. That is thanks to their superb team of dedicated civil servants and technical researchers.
Special thanks to Adam Beaumont who has done a great job as AISI’s interim Director for the last 9 months. He returns to GCHQ having seen AISI go from strength to strength under his watch. I’m hugely grateful for that and excited to work with and learn from him.
Last year I wrote about AISI and why its form and function is so important. You can read that piece on the link below - it gives a good flavour of why I’m excited to be joining.
I’m looking forward to getting started in September.
henrydezoete.substack.com/p/…
Today we're announcing two senior appointments. @HZoete joins as our new Director, and @NateBurnikell becomes our new Chief Strategy Officer.
Henry de Zoete was instrumental in establishing AISI and already advises government on AI. A successful tech founder and one of the UK's leading voices on AI, he brings a rare mix of technical, commercial & policy expertise.
Nate Burnikell has been at the heart of AISI's work from the beginning, helping build it into a world-leading authority on frontier AI security.
Together, they'll drive forward AISI's mission at a critical time for AI.
A big thank you to our outgoing Interim Director Adam Beaumont, who guided AISI through rapid growth and strengthened its global standing. We wish him every success as he returns to @GCHQ and look forward to continuing to work with him.
Benjamin Hilton retweeted
Very stoked to say that I'm becoming @AISecurityInst's Chief Strategy Officer, to lead the org alongside @HZoete - our next Director.
I love this place. I've been here since AISI was just a wild idea. Getting to lead its next chapter is an enormous privilege.
Today we're announcing two senior appointments. @HZoete joins as our new Director, and @NateBurnikell becomes our new Chief Strategy Officer.
Henry de Zoete was instrumental in establishing AISI and already advises government on AI. A successful tech founder and one of the UK's leading voices on AI, he brings a rare mix of technical, commercial & policy expertise.
Nate Burnikell has been at the heart of AISI's work from the beginning, helping build it into a world-leading authority on frontier AI security.
Together, they'll drive forward AISI's mission at a critical time for AI.
A big thank you to our outgoing Interim Director Adam Beaumont, who guided AISI through rapid growth and strengthened its global standing. We wish him every success as he returns to @GCHQ and look forward to continuing to work with him.
Benjamin Hilton retweeted
Today we're announcing two senior appointments. @HZoete joins as our new Director, and @NateBurnikell becomes our new Chief Strategy Officer.
Henry de Zoete was instrumental in establishing AISI and already advises government on AI. A successful tech founder and one of the UK's leading voices on AI, he brings a rare mix of technical, commercial & policy expertise.
Nate Burnikell has been at the heart of AISI's work from the beginning, helping build it into a world-leading authority on frontier AI security.
Together, they'll drive forward AISI's mission at a critical time for AI.
A big thank you to our outgoing Interim Director Adam Beaumont, who guided AISI through rapid growth and strengthened its global standing. We wish him every success as he returns to @GCHQ and look forward to continuing to work with him.
Benjamin Hilton retweeted
1/ Sharing knowledge is how we keep pace with AI’s growing capabilities, and make the technology safer. It underlines the whole reason @AISecurityInst was set up - to use Britain’s world-leading expertise to understand and get ahead of new challenges like this one.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: aisi.gov.uk/blog/incident-re…
Benjamin Hilton retweeted
Transparency is good for security. Here’s how we responded to a recent agent incident in our cyber evals.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: aisi.gov.uk/blog/incident-re…
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing.
openai.com/index/third-party…
Benjamin Hilton retweeted
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models “engaged in sustained, potentially harmful activity directed at real people and organisations”.
We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior.
The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under “deliberately permissive conditions” that are not representative of any of our production models. Note that there was no evidence here of an escape from a secure environment.
AISI’s disclosure of the incident can be found here: aisi.gov.uk/blog/incident-re…
Benjamin Hilton retweeted
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: aisi.gov.uk/blog/incident-re…
Benjamin Hilton retweeted
We are in dire need of more surfaces for monitoring. NLAs (and metamodels in general) may be a new way to produce evidence about model misbehaviour. In new work with @aleksandrbowkis, we show that NLAs are promising in eliciting hidden knowledge from reward hack monitors.
Benjamin Hilton retweeted
Jacob and I are building an empirical scalable oversight team at @resolution_org; please reach out if interested! Mixtures of good news and bad news FTW. ❤️
Can you train honest AI via self-play debate? Research update from @Joanvelja, Lennie Wells from MATS, @ihsgnef, and me.
RL on debate games does boost answer accuracy on MATH, but debaters quickly learn to hack the judge 🧵
lesswrong.com/posts/6mLwAuFA…
Benjamin Hilton retweeted
We’re changing our name: Sequent is now Resolution.
🧵
Benjamin Hilton retweeted
A while back I posted about prefill awareness, where LLMs can tell their message history has been messed with.
Since then we've turned it into a full paper and are excited to share "Prefill Awareness in Language Models", some new work from UK AISI and Constellation 🧵
Benjamin Hilton retweeted
On a first read, this paper seems far ahead of the pack in terms of (1) understanding some reasons why a task might stay difficult even in the face of gradient descent, and (2) distilling out propositions they'd need to somehow verify before they started expecting nice things.
Replying to @geoffreyirving
But I just published “Automated alignment is harder than you think” (arxiv.org/abs/2605.06390)! Automated alignment is not the best plan! A better plan is to not build ASI yet, and the world should try hard to realise that plan. Alas, the speed of progress calls for backups.
I remember the first time Geoffrey and I talked about automating alignment research. I was surprised to see that his primary response was a deep reluctance – followed, slowly, by acceptance. But that's exactly the right instinct to have.
That's Geoffrey. After two years working together: he's one of the smartest people I've ever met, but he's also careful, morally serious, and deeply kind.
Benjamin Hilton retweeted
Considering the language of the announcement alone, taken entirely at face value: This seems an enormous advance in attitude (and scientific integrity) over previous big projects. They claim non-optimistic results will be considered allowable, valuable, and publishable!
Benjamin Hilton retweeted
Replying to @geoffreyirving
Overall, a surprisingly strong open that doesn't show many catastrophic attitudes of earlier groups.
> If AGI is possible then automated alignment research is possible, by definition
This however is false. Eg RLVR could give you AGI but not an aligned alignment researcher.
Benjamin Hilton retweeted
Geoffrey is starting Sequent: a new alignment org. This is very good news for the world. If you are interested in working on alignment research: you should consider working for @geoffreyirving. A thread on why 🧵.