@txbraindumpi
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India App Store
Account-level information from X, not a live location or the device used for a specific post.
asking better questions
Joined December 2021
- Tweets11.5K
- Following2.9K
- Followers210
- Likes20.1K
The economics of a Neolab.
A neolab is loosely defined as a startup of AI researchers who raises a lot of money pre-production to be able to finance GPU compute to take on a large AI problem.
To buy 1000 GB300s or ~14 NVL72 racks will set you back $125-150M for 3yrs with 15-30% upfront. That’s about ~2-2.5MW.
Thats about enough to do 10^25 flops a quarter and get to a GPT-4 level model which is 1-2 OOMs off frontier for pretraining.
If you post-train on a great open source model, you have a better chance of getting to frontier. The risks are a) you need to spend millions on RL environments too and b) being lapped by another model release while being tied to a base model.
For this to payback, you need to give your customers a better and ideally cheaper inference service than a base model and serve them for long enough to recoup your large investment. Even at 50% margin on inference, to recoup $10M in training means serving ~10T tokens (!) if you price like Fable / Astra given a standard cache read / input / output split ($2/M blended). And you have to justify being better than a release like Opus 5.5 which is even cheaper. Often, you end up charging your customers a huge premium in terms of platform fees and compute fees on top of pure inference.
Meanwhile, every hour you’re not utilizing your GPUs you are burning money so you typically resell this compute back to a broker or run inference for open models / resell spot instances. At below a ~60% utilization on spot, you will still lose money.
Add to that insane cost of talent.
So what can you do with the compute?
- Not play the model game at all.
- Play an entirely different model game (Jev, World Labs) that if big labs played, would either a) cannibalize their business or b) be incrementally not significant revenue c) would cause too much distraction from the main main thing
- Acquire a proprietary data set (Peridodic Labs) in enough volume in a domain of usefulness to eclipse frontier quality. Often happens in robotics, biology, chemistry.
If you do overcome the challenge of building a model that is useful and well priced beyond big labs models, given the huge price of compute, you still need to play in an area where the revenue / compute ratio is signficant and market demand is large enough to payback your compute spend.
It is a difficult game.
For the first time, a group of researchers — including scientists from @GoogleResearch and HHMI Janelia — built the first complete brain map for a male fruit fly. Together, we mapped every single neural connection in a male fruit fly brain and central nervous system, amounting to more than 166,000 neurons.
Here’s why we did it.
ALT A colorful 3D reconstruction of a connectome shows the map of over 166,000 neurons and 125 million connections in a male fruit fly's brain and nerve cord.
ALT A fruit fly connectome map text with scattered fruit fly images. The text reads: "With it, we can compare our new map to other connectomes of the female fly brain, including our 2020 hemibrain project with partners. Having maps for both fruit fly sexes allows researchers to pinpoint exact structural brain differences and study how they drive specific biological behaviors.”
ankit retweeted
idk why his internals video doesn't get the reach on yt ... this video has an insane potential if algo work correctly
ankit retweeted
overheard in sf:
“i ran out of data i need more data so bad”
“i’m actually selling data rn what kinda data you need? coding data? healthcare data?”
“bruh i need mint mobile my plan ran out”
“oh that’s a shame. so you don’t need any coding data?”
“no i do, what kind of data do you have?”
yes, im doing it for gemma-4-e2b atp
ankit retweeted
7K+ RL environments, all in a uniform Harbor format.
> Pick a task, pick a model, pick a harness, pick a sandbox provider, and run it
ps : with OpenEnv you can directly train on these as well not just eval, there’s an Easter egg somewhere 👀
Wake up, ppl 👀
Xiaomi just open-sourced ~7K of the RL environments it used to train MiMo on HuggingFace, across domains like code, cyber, general, music and web dev
I think this is one of the biggest things to happen in frontier open-source RL environment data.
Indepth analysis and learnings coming soon 👀
a bad things about having great memory is that u learn more and fast not only applies to good information but also for bad information