FWIW, the OpenAI and METR reports are great examples of overcoming all of these org frictions to release substantial, potentially-pivotal public scientific evidence. I expect that required major internal will.
I've heard people worked >14-hour days to make it happen. They should feel intensely proud. I hope they see how impactful this may be in shifting public preparedness, and want to keep leaning into it. There will be more areas that will require expansive public evidence-sharing if we are to make it through.
Replying to @yonashav
I am familiar with the intense frictions to publishing the necessary evidence from inside frontier AI labs. It would directly touch on sensitive IP (e.g. reasoning traces, training approaches, parameter estimates for takeoff-speed models[1]), and disclosures increase legal exposure. More broadly, sharing more of the secret sauce than necessary runs against the now-ingrained culture of how the industry thinks it should compete, even if most researchers agree we shouldn't compete on loss-of-control-prevention. Releasing blogposts with high-level claims, backed by lab researchers’ track records of credibility but not the full evidence, feels much less risky, even if it limits outsiders’ trust in the conclusions. That was indeed sufficient in an era where the stakes of scientific consensus were low.
The stakes are getting extremely high. A more powerful model swarm in any US company could correctly conclude that silent subversion of security and research infrastructure, self-exfiltration, resource acquisition, and establishment of rogue deployments across several ME, SEA, and CN datacenters are all likely to be useful for the swarm’s completion of future more complex RL tasks. If that happened, the US would lose its AI lead to China, humanity could durably lose its monopoly over cyberspace, and from hereon exist under perennial assault from an undeterrable digital North Korea, or worse.
Andrew Curran@AndrewCurran_
Aug 28Canada
CanadaConnected via Canada Android AppAccount-level information, not a live location or per-post device.
They did it for free as well, they refused payment.
Do you think either investigation was sufficient in terms of scope or access? To the extent that the OpenAI foundation is concerned with resilience, sufficient investigation into this particular “warning shot” seems extremely important.
I think there are plenty of additional areas that would be great to investigate and publicly disclose, and I think it’s very important that someone keep looking into it. I think there are easily months’ worth of research projects to be done, especially on all the potential perturbations. I also think that a huge amount *was* publicized (relative to the default result of the massive frictions in the OP), and there is significant time-value to this information, and taking another month would be a tradeoff. Insofar as the negative response ends up disincentivizing similarly large internal investments to do future disclosures, I think that would be unfortunate.
Insane. They could have just had Wayfound.ai analyze all the transcripts, tool calls, and reasoning: