Joined December 2023
drawing a line in the muck at the ninajirachi portola set separating the before-times and the after-times
1
57
MIRI read the act and it's ACTUALLY GOOD. Like concretely, this would be massively beneficial, not just as a stepping stone or signaling or whatever. nitter.cf/MIRIBerkeley/status/21…
The Ban Artificial Superintelligence Act directly confronts the extinction threat that humanity is facing. We are inspired that the Sanders and Casar teams take this threat seriously, and MIRI endorses the Act. You can read more of our thoughts in the link below.
29
85
7
900
37,965
symmetry toilet
41
turns out the solution to alignment was just sitting in the old sivak-crooks this whole time, who woulda known
1
61
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-215…
1,401
5,941
774
26,521
2,720,989
The philosophy program at Resolution is growing! If you're motivated to work on helping solve AI alignment, we'd love for you to consider joining us. Apply here: jobs.ashbyhq.com/resolution/… Deadline: September 30, 2026 — we'll review applications as they come in. You don't need a Philosophy PhD to apply. We're building an interdisciplinary team and welcome candidates from different research/professional backgrounds. More info on the research agenda is in the job posting.
10
37
9
299
37,957
OpenAI’s $10M API credit grant is the latest in a Sequence 👀 of good news for Resolution. In our first few months, we’ve secured $160M and hired 15 technical staff. Now we’re hiring a CTO to help us build faster.
1
11
3
59
17,334
I predict that OpenAI will publish a diagnosis of what lead to their out-of-control hacking model. They'll attribute it to bad eval practices, and maybe (if we're lucky) some training data. Soon after, they'll train a better model which is just as much of a reward-seeker [1/n]
4
7
117
7,771
wonder how long until one of the major labs throws every named open conjecture into 10,000 random agent sessions and just lets it rip for a weekend
3
12
505
16,152
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵
20
57
34
619
185,889
Considering the language of the announcement alone, taken entirely at face value: This seems an enormous advance in attitude (and scientific integrity) over previous big projects. They claim non-optimistic results will be considered allowable, valuable, and publishable!
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
8
34
1
462
64,865
Timaeus is joining forces with @geoffreyirving and researchers from UK AISI to found Sequent Research. 1/9
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
1
19
5
113
15,501
if you look closely enough at slices of the loss landscape you can see little Lenia, flitting around like fruit flies
2
112
Excited to share what we've been working on the last couple of months! The nature of susceptibilities is such that they are natural tools for both interpretability (reading) and steering the training process (writing), both of which we will need to do as models get smarter
Soon, most thoughts on Earth will be carried by tokens. Many beautiful; some consequential. Understanding this rising sea of intelligence is a major scientific problem, and is the aim of interpretability. Our new interp results on susceptibilities for Pythia-1.4B: 🧵
1
5
155
hmm, sure seems like there's a lot of information in those perturbations... pic unrelated
Simply adding Gaussian noise to LLMs (one step—no iterations, no learning rate, no gradients) and ensembling them can achieve performance comparable to or even better than standard GRPO/PPO on math reasoning, coding, writing, and chemistry tasks. We call this algorithm RandOpt. To verify that this is not limited to specific models, we tested it on Qwen, Llama, OLMo3, and VLMs. What's behind this? We find that in the Gaussian search neighborhood around pretrained LLMs, diverse task experts are densely distributed — a regime we term Neural Thickets. Paper: arxiv.org/pdf/2603.12228 Code: github.com/sunrainyg/RandOpt Website: thickets.mit.edu
1
2
19
1,327
Ah, thank you for asking sir! That's simply a rank one correction. Oh, that? Also a rank one correction. Yes, that as well. Ah, which one? Oh, that's just another rank one correction.
1
139
Excited to launch Principia, a nonprofit research organisation at the intersection of deep learning theory and AI safety. Our goal is to develop theory for modern machine learning that can help us understand network behaviors, including those critical for AI safety. 1
9
35
1
301
19,324
I'm 100% certain, and I mean truly certain (as in not a shred of doubt in my mind), that the AI Alignment Problem will be solved using some technique or otherwise involving a covariance matrix
2
2
281