@4ldo9

Tres Comas Boom ! Sarcasm, AI & Tequila

Joined August 2022
Replying to @LiorOnAI @ylecun
@ylecun is really a boss, he called it "Le World Model" and not "Un World Model". It's like going to the Ritz an ordering un croque monsieur, on the menu it will be "Le Croque Monsieur"
1
1
381
4ldo Malkinson retweeted
If Dario wants to slow down the pace of AI, he should just move the company to Europe.
532
2,524
418
31,103
878,663
4ldo Malkinson retweeted
BREAKING: OpenAI might have stolen another major proof. In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that OpenAI may have trained Astra on conversations in which he and Gábor Kun were working on Gromov’s soficity conjecture, one of the ten problems OpenAI later announced Astra had solved. I know Andreas. We met several times early in our careers. He is an exceptional mathematician, a leading expert on sofic and hyperlinear groups, and one of the most respected scholars in the field. He has spent two decades working on this problem. If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself. And if the allegations raised by Levent Alpöge, Tristan Buckmaster, and now Andreas Thom are all substantiated, we are no longer looking at isolated incidents. We may be looking at one of the greatest intellectual scandals in the history of science. AI is not discovering new mathematics. AI is stealing human discovery.
671
4,836
788
19,821
1,800,327
CancerBench: the frontier model cancer cure benchmark. AI lab CEOs keep talking about curing cancer, so I made a benchmark. One metric: how many types of cancer has your model cured? All models are currently tied at zero. It’s time to hillclimb! cancerbench.com
138
292
48
3,205
127,053
4ldo Malkinson retweeted
Here you go: European Disapora Index. Companies valued $1B+ incl. Stripe, Snowflake, Datadog, Miro, Grammarly, UiPath, etc cc @sc_cath @lugaricano (small disclaimer: just made this today on the @dealroomco API, will refine coming days)
Il faudrait faire un CAC40 fantôme, avec les boîtes créées par des français à l'étranger. Celle-ci serait quelque part entre Renaud et Carrefour.
11
30
16
122
31,489
4ldo Malkinson retweeted
New on the geometry of motor control, introducing a dynamical systems view of motor cortex 🦾🧠 youtu.be/gAm5Athct6Y Seconds before a planned movement, motor cortex can already contain activity specific to that movement. This is the same circuit that will soon generate the commands sent to your muscles. So what is that preparatory activity doing? And why doesn't it spill out into motion?
4
29
184
9,056
4ldo Malkinson retweeted
I believe the most qualified person should get the job.
1,840
661
583
13,074
8,275,206
4ldo Malkinson retweeted
Introducing LpWM: A Case for Sparse Representations in World Models Dense Gaussian representations are a choice, not a requirement. We find that sparse representations can make latent dynamics easier to model for planning. 📄arxiv.org/abs/2608.22764 💻github.com/YilunKuang/lpworl…
14
50
6
266
44,644
4ldo Malkinson retweeted
Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:
446
1,006
546
7,042
3,538,762
4ldo Malkinson retweeted
Claude has become a language, so I built a translator. English <-> Claudish
345
1,176
255
16,784
1,140,951
4ldo Malkinson retweeted
The sense of touch is the most criminally under-explored modality in robotics. Imagine doing sleight of hand wearing thick oven mitts. That's exactly how a robot feels today if it were alive. A magnetic piece snapping into place, a paper cup peeling out of a stack, a USB negotiating its way into the port - all invisible to the camera. Learning how to feel must be a full-stack co-designed effort. We are open-sourcing a principled methodology called "T-Rex": 1. Tactile as first-class citizen of the model. Our mixture-of-transformer runs two clocks asynchronously: a slow visuomotor expert plans the motion, and a fast tactile expert refines it in real time with high-frequency corrections at 4 "touch ticks" per vision tick. Forces change faster than frames arrive, so the architecture had to as well. 2. Open data. The largest tactile dataset ever released to our knowledge: a 50-hour (~5,500 episodes) high-quality, carefully synchronized robot play corpus, collected on SOTA tactile hand hardware with 22 degrees of freedom. Available today on HuggingFace! 3. Training recipe: T-Rex extends our prior work, EgoScale. Human egocentric videos for pretraining, a diverse dose of tactile robot play for mid-training. Our experiments show this bridges contact-free pretraining to contact-rich manipulation remarkably well. Pixels are cheap and everywhere, but they run out of steam at the moment of contact. Tactile will carry the last mile. The next scaling curve will be measured in hours of touch. T-Rex is a great collaboration between NVIDIA and Berkeley: 🧵
105
130
39
908
175,163
4ldo Malkinson retweeted
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
319
1,688
859
12,174
3,412,313
4ldo Malkinson retweeted
we'll replace gradient descent within the next 6 months. it'll be the biggest change in how we do deep learning possibly ever, bigger than rnns to transformers.
60
42
11
1,020
75,882
Ok this is the best thread on X right now
gemini search is still undefeated
1
41
4ldo Malkinson retweeted
I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific. github.com/jerber/arc-code
81
130
41
1,575
397,482
4ldo Malkinson retweeted
Are robotic bugs the future of spying? 🕵🏻 Researchers at Beihang University in Beijing spent 15 years building a micro-bot that's 2cm long and moves at ultrafast speeds with no tether. It looks like a bug. Has legs controlled independently so it can move in any direction, make circles, adapt to obstacles. Can detect sounds. Small enough to fit inside a turbofan engine or be carried by a drone. Of course, the spying implications are OBVIOUS. A tiny robot that moves fast, has audio sensing, fits inside machinery, and leaves barely a trace. Theoretically deployable anywhere, ventilation systems, machinery, confined spaces. Very hard to detect. What about other applications for these kind of robot? Emergency rescue makes sense. Inspection makes sense. So let's hope these little bots will be adding value rather than spying on us :) ~~ ♻️ Join the weekly robotics newsletter, and never miss any news → ziegler.substack.com
6
15
3
148
12,940
4ldo Malkinson retweeted
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
161
292
87
3,438
476,404
4ldo Malkinson retweeted
Reward models usually get frozen after training. EvoHIL asks a simple question: what if the reward model kept learning alongside the policy? Instead of treating human feedback as a one-time annotation step, the paper lets both the reward model and the policy evolve together, creating a continuously improving RL loop.
4
15
72
4,898
4ldo Malkinson retweeted
raise $20-50M on a killer demo. hire 60-80 engineers, mostly ML, maybe 5 mechanical. spend 18 months perfecting a prototype in the lab. try to deploy in a real facility. realize the gap between lab and production is 10x your existing engineering.
7
5
91
7,523
4ldo Malkinson retweeted
τ₀-VLA is a hierarchical robot foundation model. Its high-level policy proposes candidate subtasks, a world model imagines their visual outcomes, and a value model scores task progress. Beam search and reflection then select the next subtask for a low-level VLA to execute.
1
1
1
19
3,448