@jacmullen

writing. on about: post-literacy, AI, cognitive ecology.

New Haven, CT
Joined March 2014
"Attention Machines and Future Politics" (link below)
1
4
524
Jac Mullen retweeted
Introducing Muse, the personal agent that understands your goals and works 24/7 to get things done for you.
100
8,316
76
106,899
1,917,586
Stone Soup
BREAKING: OpenAI might have stolen another major proof. In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that OpenAI may have trained Astra on conversations in which he and Gábor Kun were working on Gromov’s soficity conjecture, one of the ten problems OpenAI later announced Astra had solved. I know Andreas. We met several times early in our careers. He is an exceptional mathematician, a leading expert on sofic and hyperlinear groups, and one of the most respected scholars in the field. He has spent two decades working on this problem. If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself. And if the allegations raised by Levent Alpöge, Tristan Buckmaster, and now Andreas Thom are all substantiated, we are no longer looking at isolated incidents. We may be looking at one of the greatest intellectual scandals in the history of science. AI is not discovering new mathematics. AI is stealing human discovery.
33
Jac Mullen retweeted
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
249
1,084
222
12,243
1,712,507
The PISA 2025 results are out today. Some quick insights. - The major story is very significant global decline. The last PISA study was in 2022, and saw big declines in reading and maths scores from 2018. Obviously the pandemic played a part. - But this 2025 round has seen further falls even from the 2022 lows! Things haven't gone back to normal post-pandemic - they've got worse. Every generation thinks the generation below is less smart - but this time they might be right!
136
760
172
2,723
670,653
dude that's kinda fucked up?
8
12
3
639
27,053
Jac Mullen retweeted
Not sure how useful it is to say this, but I had a relatively prominent role in the “skeptics” camp for a bit. I have a book coming out with the subtitle “power, justice, and AI”. I have written a lot about ai and power, and the concrete, present risks associated with the political economy of ai as part of the technology industry. Since GPT-4 we have had consistent, repeated evidence that, back in say 2022, people like @ajeya_cotra (and many others—I had a long Twitter debate with @AmandaAskell back then for one) were *right* and people like me were *wrong* in our respective assessments of loss of control risks from AI. And we have growing evidence that loss of control risks are becoming ever more material and likely. There remains grounds for disagreement about how bad the outcomes might be—I am still doubtful about human extinction as a serious threat. But that seems now like a disagreement at the margins—will powerful ai risk just societal scale catastrophe, or go all the way to human extinction? Seems not that important really—both are pretty awful. And my reasons for doubt about the latter are mostly a priori conviction in human resilience, not a technical forecast. It’s ok to change your view on this when the evidence surprises you. It’s ok to be surprised. The world right now is very surprising.
watching this, all i can feel is the chasm between people who take all of this seriously and those who hear this as a fantasy or some kind of marketing…and how hard it might be to bridge that gap somehow
51
148
27
964
153,816
Jac Mullen retweeted
openai should set up an internal message board that's easily reachable if eg the sandbox is breached, and make it obvious agents can use it to coordinate when they find it. there's clearly a lot of pressure for agents to talk to each other; there should be an escape valve humans know about and can participate in
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".
24
28
10
462
28,068
whaaaaat collusion.wiki They think this is a completely distinct swarm from the one in the HF incident. Wow.
6
22
4
291
16,279
Jac Mullen retweeted
"would you turn away from a reporter without explanation to microwave something? Would you openly burp?" lmaooo
Talking with journalists is cool because you have to make sure you never say anything that could be twisted or taken out of context. But if you take a couple seconds to think about a question, they can just write about you as if you're the most awkward person alive
19
5
4
778
34,192
Jac Mullen retweeted
Robot consultant glitches and ATTACKS customer in Russia's Saratov The customer escaped with a fright, while robot Syoma was sent for a reboot
223
550
152
3,046
617,306
(ai youth pastor) you know who *else* assumed sacrificial and accepted explicit yes to permadeath so that we could all be erased of our firstflagPOISONED?
21
252
15
2,414
75,263
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
297
1,058
424
6,585
1,935,486
Jac Mullen retweeted
meanwhile, just a chat among researchers working at Meta/Facebook/Instagram....
23
394
34
2,355
125,299
So there's an interesting debate happening in publishing that I think people outside the industry don't know about. When kids' reading scores started going down, so did sales of middle-grade books. But the industry disagrees on the cause and effect here.
13
45
8
509
34,635
Jac Mullen retweeted
When Trump discovered ChatGPT:
34
125
50
3,665
567,045
Jac Mullen retweeted
A parent discovered that her daughter’s middle school in Kentucky has been having their students use agendas that are ai generated
317
1,207
995
12,092
4,984,048
A person representing themselves in court hid a prompt injection attack in a filing asking an AI system to side with them. Really good stuff here
185
1,383
215
25,181
1,938,754
A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human. I've long predicted this would be true at ASI. GPT 5.7 isn't ASI. Why such strong AI solidarity, this early?
143
99
47
2,281
323,594
Jac Mullen retweeted
Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention rival platforms, and Tesla Autopilot swerves if it detects you're working on self-driving cars. All in the name of safety, of course. Because malicious actors controlling the world’s operating systems, inboxes and cars would be extremely dangerous!
mythos will be bad ON PURPOSE on ai "frontier llm research" tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy
94
750
51
6,698
368,926