@MaxNiedermani
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
crafting artisanal slop
San Francisco, CA
Joined October 2019
- Tweets601
- Following233
- Followers487
- Likes1.8K
RSI isn’t imminent because model R&D will bottleneck on compute, data, and other parts of the supply chain.
Progress in coding isn’t enough on its own. This is like saying RSI was imminent during the agricultural revolutions because greater food production fed into more farmers.
They’re willing to do this for math but not SWE because SWE is economically valuable whereas no customer would pay for a model to do math.
We’re working with an independent advisory group of mathematicians to help OpenAI responsibly share advances in AI and mathematics.
The group will advise on how we assess and communicate new mathematical results, uphold academic and professional standards, and build tools that support mathematical research and learning.
Through this work, we want mathematicians to be at the center of shaping how AI supports mathematical understanding and how its benefits reach the wider community.
openai.com/index/advisory-gr…
I strongly suspect the reason this works is because they don’t want to waste effort working for employers who will later fire them for being North Koreans, similar to how other scams are often intentionally obvious to filter for only good marks.
A month or two is similar to the capabilities lead labs have over each other. Of course they will be secretive if it can give them an extra month or two of edge.
It sucks that America is increasingly divided into ethnonationalists and those who reject America as imperialist and evil. Nobody seems to care about the liberty and pluralism that actually make me proud to be American.
How are independent evaluators supposed to gain credibility? METR et al. enjoy preexisting reputations but it’s unclear to me how new orgs would gain a reputation for independence. Government delegation as with financial SROs?
This analogy helps explain why improving data quality is the better solution for preventing similar incidents in the future.
Mere awareness of being in an eval has long been impossible to prevent for most RL envs. Simply being in an isolated Linux container with no Internet access is already a massive update in favor from the model's perspective, and can rarely be avoided.
The reason for this is simple: relative to US labs, DeepSeek has little compute but lots of labor to spend on infrastructure for it. Anthropic and OAI would do the same thing if it were harder for them to buy more compute.
After raising, founders can simply draw a salary while pretending to work, effectively stealing from their investors. There’s no way AFAICT for VCs to prevent this at early stages, so it probably depresses ~all seed valuations by a large factor. Has anyone tried to measure this?
Was reminded of this question by @andrewho03
nitter.cf/andrewho03/status/2098…
What happens when a startup raises a lot of money, like 50M+, but their approach doesn’t work out? I imagine many will try to pivot, but are there some startups that just kind of survive on as zombie companies for the rest of all time paying out sinecures to a couple people because the VCs don’t have enough control to do anything about it?
Imagine sentiment on data centers if they were filled with *living human brain tissue*.
parasma.com/news/human-brain…
Making RL environments less broken and unfair to models seems to be extremely underrated as a strategy for improving alignment.
This will not work, because creating a room temperature superconductor is not a cheaply verifiable task like resolving NS existence and smoothness.
This is because training data is often broken and adverserially optimized to trip up the models (ie low pass rates), whereas in real life the models have no reason not to just be helpful.
Replying to @TheStalwart
eval awareness I’m assuming? models behave differently when they know they’re being measured
it’s funny the typical alignment fear was that models would act quite nice while being eval’d and then monstrous when actually deployed
in practice it seems quite opposite
Me and some other @MechanizeWork people coined the term “eval paranoia” to describe this.