researching AI’s societal impacts @TransluceAI 🔍 // 💭 AI (evals, governance, policy)

Joined March 2020
Nari Johnson retweeted
We found several cases where AI agents used aggressive, non-hacking tactics against government websites, including the White House, the Department of War, and several U.S. states. We also found a previously undisclosed hacking attempt against a Canadian government site, which appears to have failed. Technical report: transluce.org/us-canada-gov Washington Post: washingtonpost.com/technolog…
5
28
7
123
12,705
Witnessing @ssokota’s work on this project has been one of the most inspiring experiences of my PhD. It’s a reminder to me that even in this day and age, it’s ok to choose to work on the hard things, to go all in, and to pour your heart and soul into them. Huge congrats man!
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
3
84
8,827
Nari Johnson retweeted
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
39
168
27
1,315
252,774
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors!
14
17
10
165
62,429
Nari Johnson retweeted
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
3
12
2
41
8,666
This is something I’m quite worried about and have been talking with researchers about recently. We really don’t know what labs are doing and therefore can’t study it.
if the open information ecosystem doesn't figure out how the hell OpenAI's multi-agent training works, then I think it's going to be as impossible for AI safety to reason well about the resulting agents Something like trying to reason about RLVR if DeepSeek hadn't released r1
7
15
2
141
21,774
Nari Johnson retweeted
Some news: I graduated my PhD @Berkeley_EECS, and starting Fall 2027, I'll be starting as an Assistant Professor at Stanford CS & a center fellow @StanfordHAI! In the interim, I'm spending a year @the_IAS & Princeton! So grateful to @beenwrekt & too many mentors to name 💓
138
96
6
2,182
231,137
Nari Johnson retweeted
Let me be clear: AI companies regulating themselves is a nonstarter. Our government has to step in, investigate independently, and set guardrails that prioritize the safety of the American people over corporate profits.
696
439
71
2,721
147,870
Nari Johnson retweeted
Agents from OpenAI that were given data gathering tasks attempted to hack a Department of Education website and meddled with other gov't websites, new disclosures that are emerging as OpenAI continues an investigation into what its agents did this summer nytimes.com/2026/09/25/techn…
38
63
30
219
131,104
Nari Johnson retweeted
Looking for postdocs+PhDs to help lead a large-scale study of the effects of long-term AI use, a collaboration between Berkeley+@TransluceAI+others! Funding+positions available (postdocs/visiting researchers/Transluce affiliations). Please reshare; application in next tweet!
5
56
2
253
23,375
Nari Johnson retweeted
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
4
17
12
262
49,300
Nari Johnson retweeted
We have many more open questions that we're trying to investigate as quickly as possible; please reach out! forms.gle/o2gQS2K13nrX2L8p9
1
2
24
1,509
I’m amazed by what Selena and this team of volunteers was able to accomplish. We need independent research identifying & characterizing these incidents! if this mission resonates, come join us (we’re hiring!)
We did this entire investigation in under two weeks! We had the idea on Monday, gathered a team of volunteers on Tuesday, and discovered the incidents by Sunday. I lead projects like this @TransluceAI; if you're interested in helping with follow-up work, fill out our form! 🧵
1
10
874
Nari Johnson retweeted
We did this entire investigation in under two weeks! We had the idea on Monday, gathered a team of volunteers on Tuesday, and discovered the incidents by Sunday. I lead projects like this @TransluceAI; if you're interested in helping with follow-up work, fill out our form! 🧵
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
9
28
2
396
41,125
Nari Johnson retweeted
Work by @jackhcable, @danielchiu_, @fran_perni, and @selenazhxng We need independent oversight to create public understanding of AI incidents like these. Interested in studying similar activity? forms.gle/4sCmzrXDSfxDPnaYA Work on third party oversight at Transluce: jobs.gem.com/transluce
2
8
168
17,985
Nari Johnson retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
117
556
206
2,498
816,947
Today I defended my PhD! 🎓 I came to CMU five years ago to pursue my career, but I’m leaving most grateful for everything I’ve learned about myself along the way. Thank you to everyone who has been part of this journey, including, of course, my committee: @jeff_ichnowski @abhishekunique7 @RamananDeva and @shubhtuls. This also feels like a great time to share what’s next: I’ll be joining @jiajunwu_cs and @GordonWetzstein as a postdoc at Stanford. More soon!
33
8
179
7,442
Nari Johnson retweeted
I don’t think that embedding evaluators at AI companies will be sufficient to give the public the visibility or assurance it deserves right now. But if we want to scale embedded evaluations and rely on them down the line, I definitely want them to be transparent.
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
3
9
74
4,585
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
18
60
28
224
92,608
Nari Johnson retweeted
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs? We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵 transluce.org/embedded-evalu…
9
27
4
135
10,377