@MangoSweet78i
iAccount based inCanada
About this account
- Account based in
- Canada
- Connected via
- Canada Android App
Account-level information from X, not a live location or the device used for a specific post.
gay leech of admiration I do training @ 🐬 & build gunpla. banner: @N8Programs
Québec
Joined June 2024
- Tweets14.4K
- Following1.7K
- Followers935
- Likes89K
Pinned Tweet
A ton of work went into creating the env & model, Dozens upon Dozens of runs, Love how it turned out and there's so much more coming soon. Can't wait to share it all. Go read the blog post!!
Dolphin X1 Trinity Nano is now live on @huggingface
Our smallest decensored model yet - 6B MoE with 1B active parameters trained using only online RL
Huge thanks to @TargonCompute for providing an 8xB200 node, @PrimeIntellect for hosted RL, and @arcee_ai for the Trinity series
Moving is fun if u do it with a friend we just sang 'we are charlie kirk' the entire fucking time, New neighbours already prob hate our asses.
🥭 retweeted
I used 15 harnesses to create Hello World with Fable 5.1 @ max and all results were the same
Thus I conclude harnesses don’t matter
🥭 retweeted
A Domain Expansion is achieved through an overwhelming sense of self and represents the ultimate manifestation of a sorcerer’s ego.
Sukuna’s divine feat of achieving an Open Barrier is representative of his ability to project his sense of self onto the world.
Megumi’s incomplete domain represents his half-hearted sense of ego, his overwhelming sense of self isnt complete because he’s a selfless person, and its shown via his Domain buffing his Shikigami and technique. He has an emotional connection to his Shikigami and fights WITH them.
Yuji has no overwhelming sense of self. His Domain was achieved through an overwhelming sense of selflessness. His Innate Domain manifests as his home town, and requires manual activation of his sure hit because he believes in giving others a chance at redemption. He is compassionate and we literally get shown a buddha statue and his mudra is that of a boddhisatva who delays his own enlightenment to guide others to enlightenment.
His Domain doesnt need a name because it goes against the very nature of what a Domain is.
And this is expanded on in Modulo with his self-isolation and “old soldiers never die, they fade away”. He has no overwhelming sense of self, he echoes Gojo’s wish to be left behind because thats the role he gave himself. Its also why he vows to fully train the next generation, track down every HR lineage and turn himself into a Curses Object. His role is to guide others.
Thus his Domain doesnt require a name.
Replying to @perksverse
Oh, god it's so retarded. We have no proper working theory on either LLM or human congnition, but we do have two narcissists in Academia that try to prove something deep relating to both.
Throw it straight to garbage.
🥭 retweeted
Is academia just worthless performance art at this point?
Oxford researchers argue that LLMs can never invent anything.
It is mathematically impossible.
They published a paper called “Theory Is All You Need" and it argues against the claim that computational models can generate genuine novelty or new knowledge.
They analyzed the limits of generative ai, and the results are a brutal reality check for the idea that ai will replace human decision making under uncertainty.
Here is why AI is stuck and human cognition wins:
backward-looking vs forward-looking.. llms are probability machines that look backward at existing data. human cognition is forward-looking and capable of generating genuine novelty. human cognition operates theoretically "top-down" rather than "bottom-up" from data.
the "data-belief asymmetry".. the researchers use the invention of "heavier-than-air flight" to illustrate this concept. an ai relies on data-based prediction, which is largely imitative. humans, however, use theory-based causal logic that allows them to hold beliefs that go beyond existing data.
the intervention gap.. humans don't just process information; we use theory to practically "intervene" in the world. we engage in directed experimentation to generate entirely new data. ai-based models are theory-free and place primacy on existing data and prediction.
tldr?
AI uses a probability-based approach to knowledge and ia largely imitative. It can process data and make predictions, but human cognition relies on theory-based causal reasoning.
The decades-old analogy comparing human minds and computers to mere "input-output" devices is fundamentally flawed.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: mimo.xiaomi.com/rl/
Complaining about like the 30 minutes steps I'm getting right now while MiMo is getting like 5 hour step times.
Fish+Wezterm looking really good.
Regardless of I like Wezterm, I'm definitely sticking with Fish from now on.
🥭 retweeted
Thanks to @maria_rcks and @theo, all your Codex threads stop running when you hit your usage limit
Thanks both!