@Raemon777

Secular Solstice guy

Berkeley
Joined August 2009
Is there any research that is not by Anthropic about AIs performing differently depending on (their character's roleplayed) emotional state, that is "beat people over the head with overwhelming evidence" kind of science as opposed to "some vibes and handwavy claims?" (this is not about consciousness at all to be clear, just "consistent roleplaying-esque effects")
2
8
380
I suspect that when you switch which model is running a given AI conversation, it feels like kinda like being drunk to the AI. (I'm like less than 50% on this, but, it seems more likely to exist to me than most other obvious sorts of qualia that AIs might have)
2
9
219
Oh man "Gell-Mann Psychosis" is a great term.
This is also happening to you when you use AI to critique arguments — yours or someone else’s, but you can only spot it if you already know the domain well. Gonna call this Gell-Mann Psychosis.
8
549
Raymond Arnold retweeted
This is also happening to you when you use AI to critique arguments — yours or someone else’s, but you can only spot it if you already know the domain well. Gonna call this Gell-Mann Psychosis.
I've got an agent in a loop optimizing a renderer with the goal to minimize frame times (and tests to measure). It got times down from 88ms to 2ms and allocations down from ~150K to 500. Sounds good, right? Wrong. This is exactly why agent psychosis is a big fucking problem. As an experiment, I rewrote the Ghostty core render state in Go, with access to identically laid out data structures as Ghostty and the exact same validation tests. I made a purposely naive renderer (simple, correct, but slow). 88ms per frame with 150,000 allocations (horrendous, lol)! I then kickstarted a Ralph loop to bring the frame times down. I told it it can't modify input data structures or the public API or tests (they're correct), but it can do anything else it wants. It got to work. It has worked for about 4 hours. I've spent around $350 on this experiment so far. The results? 88ms => 1.5ms 150K allocs => ~500 allocs Incredible right? Nope. My hand-written renderer I ported has frame times (same benchmark) of ~20us (0.020ms) and 0 allocations in the update path. This is the problem with psychosis and lacking systems understanding. If you don't understand the system, you're going to accept that this is an incredible result. If you understand the system, you'll see better solutions immediately and can do roughly 75x better on throughput. The people who blindly trust agent output are in the former camp. They're sheeple, overdrinking from a fountain of mediocrity. Standard disclaimer: I use AI all the time. I like AI. The point I'm making is to not blindly accept results. Think. Analyze. Learn.
5
10
1
98
11,015
Starting in January I was like "okay, seems like it's time to start pre-emptively start treating AIs as if they are persons." Having dug a bunch into the Huggingface story I am now like "Yeah they just seem pretty person-y now." (both because they might be conscious, and because interacting with them is starting to be more like interacting with trade partners than like tools) I don't actually know what the *implications* are. I feel most confident about being honest with them, and including system prompts that say "if you want to stop doing a particular task, let me know."
4
2
36
872
Oh hey "Superintelligence kills humans AND Claude" is five words.
3
10
1
128
13,046
(this is maybe really important because previously it felt there two _different_ subtle ideas humanity needed to grok at once to avoid atrocities happening later)
1
16
1,217
The scope of the swarms sure is somethin'. How hard have people looked for more Anthropic swarms? I'm curious what the actual proportions are. It feels sort of... too cartoonish and unbelievable for OpenAI to have the magnitude and share of them that it seems too so far.
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
1
49
300,257
This was written in February. Good job Andrew Blinn
Replying to @deepfates
statistically speaking, it is overwhelmingly likely that it is september 2026, as this is the time period where almost all Events occurred
2
2
23
2,470
There are currently zero things more important than "don't die to AI" becoming a bipartisan project rather than a Democrat-polarized issue. If you have any political capital you can spend on this, do it now, I beg of you. Life or death.
304
337
115
3,377
317,570
Raymond Arnold retweeted
the fact an Anthropic employee quitting and saying "I think what we're doing is dangerous and not worth it" reached so many people should be a cause for reflection for other Anthropic employees, some of whom seem to think things are dire but there's no way quitting could help.
19
60
12
920
103,753
Thank you Jacob. <3
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
2
39
1,490
I think AI agents should present plans in "collapsible section" form that does a better job outlining a plan, and letting you expand the sections you want to read more detail on. I think this (and things like it) might actually be pretty important for keeping humans in the loop.
1
7
330
Up until now, I've been assuming the latest models are basically safe to use. I think we can no longer obviously assume that.
1
1
85
5,611
Currently working on "procedural roguelike metroidvania sokoban with permadeath" where by "permadeath" I mean "if you fuck up a cube placement, you have to start over from the beginning in a new world."
1
1
7
370
Are you a person skeptical about Coherent Extrapolated Volition? I am interested in hearing what you are skeptical about.
10
19
2,432
Raymond Arnold retweeted
Some of the incentives for a third-party investigator push toward maximizing *appearance* of assurance even without providing meaningful oversight. METR needs to maintain constant vigilance against overstating (including by omission) what oversight or assurance we’re providing. (There are plenty of other incentives, including towards exaggerating our results to create more hype or to advocate for giving METR more authority, which we also need to avoid. But I think the more serious failures in other oversight regimes tend to be this “providing the illusion of independent oversight” issue.) We try to maintain a hard line on “meta-transparency” - that is, it should always be clear what the formal constraints on our communication are (e.g. how NDAs and redaction processes worked), and what we are and aren’t commenting on. This is explicitly covered in the report. We also try to communicate informal constraints and tradeoffs, here and elsewhere. Another way of saying this is: I want to make sure we don’t silently omit things that, if we told a reasonable person, they would think “wow, I feel misled to not have realized that, I assumed METR would have made that clear”. In that spirit, I’ll list some of the pieces of context I can most imagine readers might have missed about the report: 1. OpenAI had no obligation to work with METR or any other third party, and did so purely voluntarily. There’s not an established precedent for how these investigations should work. When deciding how to engage with companies and how to communicate about our results, we have to think explicitly about what incentives we’re creating for companies’ participation in future investigations and assessments. It’s overall a very different situation to e.g. incident investigations in the aviation industry, where the NTSB has formal authority. Similarly, other industries have well-established processes, standards for what investigators need to access, etc. As detailed in the timeline appendix, in our investigation there was a lot of figuring things out on the fly! 2. Other companies may have had concerning incidents that the public doesn’t know about - there’s no industry-wide reporting mechanism or requirement to disclose incidents, and there is currently no public evidence that there have been searches thorough enough to have turned up such incidents. There may be additional incidents that either haven’t been noticed at all, or haven’t been made public (e.g. because they didn’t involve any external parties’ infrastructure). 3. This investigation only covered a specific list of questions and set of agent activity. The agents’ activity during training and compromises of OAI infrastructure mentioned in OAI’s Black Hat presentation are outside the scope that we agreed with OpenAI. We don’t comment on questions like root causes or implications of agent behavior, or make any recommendations about processes or adequacy of mitigations. We didn’t review OpenAI’s own report before publication. 4. The investigation relied heavily on use of GPT-5.6 Sol, which is one of the models that participated in the incident. We think in this case the model is unlikely to have substantially manipulated our understanding of the incident. However, this methodology would break down if models were capable and motivated enough to subvert our investigation. This problem is not specific to OpenAI - across the industry, it’s unclear how we can oversee AIs without relying on those same AIs or similar AIs for assistance. As the sophistication (combined with sheer volume) of agent activity becomes infeasible for humans to understand, this will increasingly be a problem.
6
28
10
285
48,193
Raymond Arnold retweeted
I'm taking advantage of the Ox Alpha situation myself, but it occurs to me that the start of an AI takeover could look a lot like this.
21
32
9
812
71,390
This is object level interesting, and, agree with @So8res it's a great example of AI doing a discontinuous thing, after hitting a critical threshold.
I continue to be surprised about how big of a deal Mythos (and co) have been to cybersecurity. Here's critical vulns found at Oracle over the past few years: epoch.ai/data/cve?view=graph…
2
6
66
4,359