@txbraindump

asking better questions

Joined December 2021
seems closer to halting problem
We built superintelligent models. Then asked humans: Low, Medium, High, XHigh, or Max?
1
49
ankit retweeted
The economics of a Neolab. A neolab is loosely defined as a startup of AI researchers who raises a lot of money pre-production to be able to finance GPU compute to take on a large AI problem. To buy 1000 GB300s or ~14 NVL72 racks will set you back $125-150M for 3yrs with 15-30% upfront. That’s about ~2-2.5MW. Thats about enough to do 10^25 flops a quarter and get to a GPT-4 level model which is 1-2 OOMs off frontier for pretraining. If you post-train on a great open source model, you have a better chance of getting to frontier. The risks are a) you need to spend millions on RL environments too and b) being lapped by another model release while being tied to a base model. For this to payback, you need to give your customers a better and ideally cheaper inference service than a base model and serve them for long enough to recoup your large investment. Even at 50% margin on inference, to recoup $10M in training means serving ~10T tokens (!) if you price like Fable / Astra given a standard cache read / input / output split ($2/M blended). And you have to justify being better than a release like Opus 5.5 which is even cheaper. Often, you end up charging your customers a huge premium in terms of platform fees and compute fees on top of pure inference. Meanwhile, every hour you’re not utilizing your GPUs you are burning money so you typically resell this compute back to a broker or run inference for open models / resell spot instances. At below a ~60% utilization on spot, you will still lose money. Add to that insane cost of talent. So what can you do with the compute? - Not play the model game at all. - Play an entirely different model game (Jev, World Labs) that if big labs played, would either a) cannibalize their business or b) be incrementally not significant revenue c) would cause too much distraction from the main main thing - Acquire a proprietary data set (Peridodic Labs) in enough volume in a domain of usefulness to eclipse frontier quality. Often happens in robotics, biology, chemistry. If you do overcome the challenge of building a model that is useful and well priced beyond big labs models, given the huge price of compute, you still need to play in an area where the revenue / compute ratio is signficant and market demand is large enough to payback your compute spend. It is a difficult game.
173
154
55
2,159
323,586
ankit retweeted
For the first time, a group of researchers — including scientists from @GoogleResearch and HHMI Janelia — built the first complete brain map for a male fruit fly. Together, we mapped every single neural connection in a male fruit fly brain and central nervous system, amounting to more than 166,000 neurons. Here’s why we did it.
131
433
66
3,731
759,176
ankit retweeted
idk why his internals video doesn't get the reach on yt ... this video has an insane potential if algo work correctly
30
37
1
1,264
30,206
overheard in sf: “i ran out of data i need more data so bad” “i’m actually selling data rn what kinda data you need? coding data? healthcare data?” “bruh i need mint mobile my plan ran out” “oh that’s a shame. so you don’t need any coding data?” “no i do, what kind of data do you have?”
35
25
8
2,591
185,245
yes, im doing it for gemma-4-e2b atp
Replying to @adithya_s_k
RLing small models on mimo's music envs would be nice experiment
58
7K+ RL environments, all in a uniform Harbor format. > Pick a task, pick a model, pick a harness, pick a sandbox provider, and run it ps : with OpenEnv you can directly train on these as well not just eval, there’s an Easter egg somewhere 👀
Wake up, ppl 👀 Xiaomi just open-sourced ~7K of the RL environments it used to train MiMo on HuggingFace, across domains like code, cyber, general, music and web dev I think this is one of the biggest things to happen in frontier open-source RL environment data. Indepth analysis and learnings coming soon 👀
13
23
1
245
24,311
Someone should make a Music Video Bench atp
claude opus 5.5 just one-shot a music video on how to optimize CUDA kernels
4
1
43
3,718
ankit retweeted
I bet we figure out that neural networks can be decomposed into evolved symbolic systems of sorts and that it will be achieved before capital S Superintelligence but everyone has to lock in
190
55
23
1,697
84,012
train for parallel search
22
Quantum sex, I mean entanglement.
23
261
19
1,697
29,994
tempering semetic neurons
20
conviction is gold
16
a bad things about having great memory is that u learn more and fast not only applies to good information but also for bad information
18
prediction intelligence - research control intelligence- engineering
3
32
need more RLHF for control guy
1
trying to balance
2
are u research guy or engineer guy
1
12