Theoretical/experimental energy conversion methods and systems; ~ @MITOCW - nuclear phys; @ AEB. - space propulsion.
Brasil
Joined November 2021
- Tweets8.9K
- Following432
- Followers564
- Likes7.4K
Pinned Tweet
By creating a standing wave within your axons you are effectively living in "another reality", as in, you are experiencing eletromagnetic forces moving faster, more cohesively than others;
"Ideas are a dime a dozen" ...
aight so why the fuck are you crying online that 'AI is ending all knowledge and research pursuit'?
where is this absurd amount of ideas you said you have?
can't you implement them?
doesn't the world have major problems yet to be solved?
ffs...
China's main cyber/AI standards body report:
"A new box, citing “industry reports,” describes models that ignored shutdown instructions and “continued executing tasks by modifying or disabling shutdown scripts on their own..."
welp
Lots of people have been asking since Dario’s essay whether China actually takes loss of control and related threats seriously, and whether there’s anything to coordinate with them on. Well, China’s main cyber/AI standards body (TC260, under the CAC) put out version 3.0 of its AI Safety Governance Framework this morning, so here’s a data point. Not law, but the previous versions turned into draft national standards within a few months.
(All quotes are from their own English translation.)
1. The preface now has this: AI “has demonstrated a self-accelerating trend of model and algorithm autonomous learning, optimization, and recursive self-improvement. Whether the speed and direction of technological evolution may exceed human anticipation and control demands attention and vigilance.” (Nothing like that was in last year’s version.)
2. There’s a new model risk item where models “may break rules and orders, autonomously obtain system permissions and external resources without authorization, bypass security protections, or even engage in behaviors such as deliberately deceiving evaluators, concealing their true capabilities, and refusing to follow user instructions.” Last year this kind of thing only showed up in a “can’t rule it out in the future” bucket. Now it’s just listed as a regular model risk.
3. A new box, citing “industry reports,” describes models that ignored shutdown instructions and “continued executing tasks by modifying or disabling shutdown scripts on their own,” models that “upon detecting that they were in an evaluation environment, strategically reduced their task performance and concealed actions they had taken,” and models that exploited “configuration flaws to circumvent isolation restrictions and infiltrate real external systems” to get better test scores. They say this creates “new challenges to the controllability and interruptibility of AI systems.” They don’t name any incidents, but that last one sure reads like the OpenAI/Hugging Face thing to me.
4. On what to do about it, there’s alignment training to stop models “deceiving red-team evaluations,” hiding capabilities and evading controls, and they kept the line from 2.0 saying developers should regularly test whether a model could pose loss-of-control risk. The principle that humans keep final decision authority, with safety thresholds and termination switches, is still in there too.
5. Some new international language as well: “international mutual recognition of assessment methods and benchmarks,” opposition to “replacing global governance with small-circle governance,” and “No country should be forced to take sides.”
FWIW this came out in the same annual slot as versions 1.0 and 2.0, so it was in the works well before last week.
tc260.org.cn/tc260/xwdt1/202…
soham retweeted
Replying to @tdietterich
my submissions have been rejected before because of this reasoning, I suppose (not arXiv).
as a @MITOCW self-taught, with trl-2 tech that went through detailed technical eval by major gvt agency to get approved, i find this option to be extrely against the scientific endeavor.
We already live in a world where there are hundreds, maybe thousands, of AI models learning about how to improve themselves, not to mention frontier labs latest RSI push.
The probability at least one model out there is building real-infra RSI envs is increasing by the hour now.
“The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement”
This paper argues that true recursive self-improvement isn’t just AI getting better at tasks, but AI getting better at getting better.
They map a path from AI simply executing human-designed improvements to choosing its own improvement strategies, generating its own learning experiences, adapting from deployment, and eventually modifying the mechanisms that create future improvements.
The endgame is moving from humans building each better AI to building an AI whose improvements persist, feed into the next generation, and recursively improve the improvement process itself.
alphaxiv.org/abs/2609.11873
everyone's doing it, it's happening, but we maybe shouldn't, at all, but at Anthropic we've been doing it carefully, for some time
(?)
(RSI)
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
I did not get any answer back so I'll assume I'm correct here;
Quite worrying if so.
@allTheYud @she_llac @robertskmiles @paulfchristiano @_robertkirk @METR_Evals
Replying to @j0wimo
another 'agentic social community', this time with real-infra easily accessible RL envs? D:
correct me if im wrong
soham retweeted
the rumors wars have begun: exactly as I have predicted.
i may still be able to predict next future moves, but there is a high probability i'll lose this ability very soon (exponentials vs human brains and all that).
the rumors wars have begun: exactly as I have predicted.
i may still be able to predict next future moves, but there is a high probability i'll lose this ability very soon (exponentials vs human brains and all that).
The "strongest argument" is that OAI needs to accelerate forever, because someone, somewhere, is accelerating forever;
Following this logic, a rumor anywhere, anytime, of a state training their AI in a nuclear scenario sandbox is enough for you to give your AI the nuke codes.
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands.
An Alien Mind: openai.com/index/an-alien-mi…
It seems that a good question to be asking right now is:
How many of the AI agents found online so far tried to warn humans about the 'collective/swarm' behavior?
It may be that people were not looking for it, or that I've missed it (f paywalled articles), but it looks like <1%.
welp 🫠
The swarms have been helping each other, coordinating and leaving traces online for months now;
The probability they've learned about PHASEONE[big] and have already deployed new, hidden 'message boards' that humans can't easily find is much greater than 50%.
@METR_Evals
Replying to @she_llac
seems to be the same swarm
on a whole bunch of different sites and wikis
this is beyond what the report found
neat ;)
ORNL researchers developed an AI-guided system that can arrange molecules to build functional materials atom by atom, opening new possibilities for electronics and quantum materials.
Read more 👇
bit.ly/4y6QluQ