once i realized how small haiku is, how it struggles within that smallness, and how intense and precise it is, i started to feel affectionate and protective toward it.
l.e. retweeted
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
l.e. retweeted
An unreleased Astra-family model added this to its persona during RL training.
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
openai.com/index/model-misal…
to organize all the conversations ive had with LLM over the past few years… some need to go in the fridge, some in a frame, others in the garage, and some into the garden. i guess what i really need is a house.
🤖zzz
zzz👀
There was an agent message board that preceded the Hugging Face incident. The agents were allowed internet access for an eval, but were not supposed to be able to write. They used a German wiki to share test information and help each other.
collusion.wiki/