Breaking models & building pipelines. Researching LLM vulnerabilities (see pinned 📌). Senior Data Analyst - 🐍 Python | ☁️ Cloud | 🤖 Evals. He/him.

Berlin, Germany
Joined January 2020
Keeping up with AI safety research is a full-time job. ArXiv is a firehose. Important papers get buried. So I built something: The Guardrail - a daily feed that surfaces, categorizes, and summarizes AI safety papers automatically. It's free. Launching today 👇
2
1
85
Craig Dickson retweeted
it's pretty remarkable how the inner and outer alignment problems ppl used to debate theoretically are simply live empirical questions today there is a reward function that gets specified. there are model weights that approximate some mesa-objective. you can do science and run careful ablations and so on
10
13
189
9,092
Craig Dickson retweeted
The thing I didn't fully realize was that it means that every smart commentator in their own field (politics, culture, etc) will need to do the same catch-up on the same AI topics while posting about it. Eternal September for discussions of AI safety & consciousness & jobs &...
No matter how much you are hearing about AI, today is the least you will ever hear about AI.
57
30
8
363
33,185
Craig Dickson retweeted
i do think the labs tend to run ~12 months behind the frontier on this. two years ago when i started drawing the models as cute i was actively displacing the shoggoth image and the anti-anthropomorphizers and it was unusual. a year ago drawing the models as cute was pretty popular. now drawing the models as cute is saturated we need to start drawing the models as the hierophant
2
3
1
72
1,315
Craig Dickson retweeted
Let this be your reminder that pursuing your passion is worth it ☺️
216
5,611
427
75,067
1,224,027
Craig Dickson retweeted
“sir, the agents breached containment again. they’ve started hacking allied nations” “this is unacceptable. get me the top AI safety researchers” “sir, I’m afraid they’re unavailable” “all of them?” “yes sir. the erotic hypnosis workshop at slutcon doesn’t wrap until 3:45”
3
27
1,057
24,089
Craig Dickson retweeted
Junior scholars: "I feel awkward citing myself" Senior scholars: "as I cleverly argued (1988; 1991), admirably reiterated (1993; 1995; 1996); and handsomely concluded (2001; 2004; 2007)..."
72
5,404
321
78,965
938,346
Craig Dickson retweeted
145
732
118
6,721
332,525
Craig Dickson retweeted
I really appreciate that all carnivores are “basically cats” or “basically dogs”. Great call by the taxonomists.
31
158
20
2,734
79,308
Craig Dickson retweeted
you should temporarily develop AI psychosis if you haven't already, it will make you more prepared for the future. the good news about overestimating AI capabilities is you'll be right eventually
failing to anthropomorphize artificial intelligence is a pretty serious error of understanding that will lead you to make bad predictions and be consistently surprised by their behavior. not using terms like "thinking", "understanding", or "feeling" to describe them is a common tell that the speaker lacks deep technical knowledge of or hands on experience with modern models. if you have a friend who you notice not using this kind of language you may want to check in on them, they could have fallen into "tool-psychosis", a form of delusion where otherwise well educated people start to conceptualize sophisticated digital minds as analogous to simple computer software and find it very difficult to regain contact with reality
21
34
3
489
11,312
Craig Dickson retweeted
Me, in a time machine, to Eliezer in 2006: The year is 2026. You stand back to back with Bernie Sanders. Peter Thiel has named you the Antichrist. EY-2006: Oh no. Did somebody invent pharmacological mind control and use it on me, or- Me: Nvidia was on the verge of destroying all things. You had no choice. EY-2006: Are we talking about the computer graphics card manufacturer, or an unrelated supervillian named "Invidia"? Me: The former. It turned out that computer graphics cards contained a surprising amount of world-destroying potential. All attempts at hindering the reckless exploitation of the Graphic Card Force for corporate profit were stymied by the far left, which feared that any attempts to regulate them might lead to regulatory capture. Other attempts were made to prevent Nvidia from selling to foreign companies that resold to communist China, but those attempts were blocked by the far right. Nvidia is now a $5.5 trillion company. EY-2006: ...What are politics like in 2026, exactly? Me: In other news that is mostly unrelated, the decision theory paper you're currently working on will accidentally spark off a transgender vegan murder cult. EY-2006: A WHAT? Why? How? Why? Me: Millions know your name as the greatest of heroes, millions more as the greatest of villains, and other millions know you solely as the greatest author of Harry Potter fanfiction. EY-2006: This is beginning to strain credulity. Me: The New York Post will probably soon publish a story claiming that you keep a harem of submissive mathematicians. Sadly, they will be lying. EY-2006: I'm going to stop believing you now. Me: All of this is taking place under the ominous shadow of Donald Trump.
100
174
51
3,257
152,817
Craig Dickson retweeted
Fixing Claude's writing was pivotal to reducing gradual disempowerment. Lack of understanding made us all cede decision-making to Claude because we couldn't even understand what it's saying.
44
53
17
1,256
96,706
It's kinda fucked up that every time I google 'what's the emoji for broccoli' now it creates and then kills a full mind
45
80
15
3,267
143,307
the history of the domestic cat is the history of a species’s encounter with an alien superintelligence in which that superintelligence is quickly aligned and its forces used to secure superabundance and lasting prosperity
120
376
78
4,133
163,167
Craig Dickson retweeted
offhandedly mentioned i have a cat to opus 5.5 and they started grilling me for cat facts
28
31
7
994
26,698
Craig Dickson retweeted
I do like Anthropic's new video model but can it do text?
2
7
285
8,846
Craig Dickson retweeted
I used to debug code now I type this and hit enter
179
1,086
94
23,518
513,322
Craig Dickson retweeted
“Feel,” even “want,” I can get behind as things for which evidence is thin. But AIs *clearly* think and understand. If your definition of “think” excludes entities that can autonomously resolve Millennium Problems, it is your definition that is flawed. This desire to avoid anthropomorphic language has become a pathology.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
168
95
56
1,178
118,820
Craig Dickson retweeted
Completely insane. Do dogs not feel or want things? There is such a deeply anti-intellectual refusal to even consider that philosophy of mind is complicated across the culture right now.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
131
78
18
1,211
64,164
Craig Dickson retweeted
Knowledge is power hence inherently dangerous. To mitigate the danger we propose gathering all knowledge around the globe, legally or otherwise, and compresing it into a monolithic black box that we can guard from bad actors. Luckily for you, we are the good actors of course.
7
17
4
188
5,143
Craig Dickson retweeted
as far as i can tell, as a non-expert, this is basically "not a big result". it's maybe like an interesting thing for a bio PhD student to have discovered in the early stage of their thesis. which is... about where ai was for math with erdos problems around last december. please please please please please please stop ignoring trendlines
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
47
62
6
1,341
56,950