growth @nebiustf @nebiusai // ex @Scaleway // from silicon to token, inference and anything in between. Views are my own - not financial advice

Paris / Barcelona
Joined January 2022
Name a better hackathon venue…i’ll wait @hackbarna
3
3
36
3,389
Typesafe will be a $100b company in 10 years
Instinct will be a $ 100b company in 10 years.
6
3
66
17,879
dylan ツ retweeted
we live in the dumbest timeline, maybe we do deserve to all die at the hands of intelligence
12
12
1
296
17,440
$20k grand prize is insane nebiusglobalaihackathon.devp…
Show us what you’re building 👀 Join the NVIDIA × Nebius Global AI Hackathon powering agents, ai apps, coding tools, physical ai... $50K+ in prizes. Bring your best projects. 42 days left! link below
2
4
72
11,338
dylan ツ retweeted
The labs committing to gigawatts are still looking for capacity in tens of megawatts. That is a more interesting signal for Nebius than another announcement about the eventual size of the AI market. CNBC reports that Anthropic has explored 20–30 MW deployments in the UK and Nordics, while OpenAI has explored smaller deployments in the Nordics. My read: customers are buying two things when they buy compute. Processing capacity, and the ability to start using it at a particular time. Inference makes that second dimension especially valuable. A large training run needs many GPUs communicating closely. An inference fleet can serve separate requests through independent copies of a model. Each copy may still require substantial, tightly connected infrastructure, but separate copies can operate in different locations. That changes which sites are economically useful. A pocket of available power that cannot support a giant training cluster may still support meaningful production inference. An operator with several suitable locations has more opportunities to match a customer’s workload and deadline. This is the strongest argument for the distributed part of Nebius’s strategy. Its four announced UK deployments are expected to reach 65 MW combined in 2027. They sit alongside much larger projects, including the planned 1.2 GW Pennsylvania campus. Together, these give Nebius several ways to add capacity as demand develops. There is already a commercial signal behind this. For a customer constrained by compute, waiting can mean delayed launches, tighter usage limits and demand it cannot serve. A lower future infrastructure price has to be weighed against those costs. The advantage still has to be earned. A 20–30 MW allocation can sit inside a larger campus, so this reporting does not establish that customers prefer smaller buildings. And distributing capacity creates its own problems: duplicated model copies, uneven demand, more operational work and limits on where customer data can go. Nebius has to keep those sites well utilized and deliver consistent performance. Otherwise, a broader footprint simply becomes a more complicated one. But the strategic logic is strong: regional deployments create additional opportunities to serve demand while larger projects progress. For us, the opportunity is to make more of the world’s available power useful to customers sooner. For those customers, the cost of compute includes the cost of waiting.
OPENAI & ANTHROPIC HUNT FOR 20-30 MW AI CAPACITY IN UK, NORDICS OpenAI and Anthropic are looking at smaller compute deployments across the U.K. and Nordics, with similar discussions also taking place in the U.S., per CNBC. Anthropic has explored 20-30 MW deals in the U.K. and Nordic region, while OpenAI has been looking at opportunities in the Nordics. The move would complement their much larger multi-hundred-MW and gigawatt projects by giving both labs faster access to powered capacity that can come online sooner. These smaller clusters are particularly useful for inference, where AI workloads can be spread across multiple sites instead of relying on one massive tightly connected training facility. That shift is becoming more important as more compute moves from training models to serving them in production. JLL expects inference to overtake training as a share of data center capacity in 2027 and reach 37% of global workloads by 2030. Anthropic already has a roughly $45B deal with Nscale for about 460 MW in West Virginia, while OpenAI has committed to several multi-GW Stargate projects across the U.S. Source: CNBC
2
4
37
4,117
dylan ツ retweeted
almost every high performance motor on earth turns on a sintered rare earth magnet: airpods Phone haptics EV drive units robot wrists missile fins cmpressor fans in a hall full of GPUs the bottleneck in three numbers: - China mines about 60% of the magnet rare earths - it refines about 91% - it makes about 94% of the sintered permanent magnets having deposits might be a comforting idea, it's still just a small part of the strategy. Ore is not the product. The product is basically a brick that keeps its field when the motor gets hot. Building that brick means separating periodic table twins, alloying them, powdering, sintering, machining, magnetising. China built that middle when the west treated it like a dirty job nobody wanted at home. so who owns separation, and a sinter line a customer will actually trust? A few names sit on those steps: - $MP : Fort Worth metal + magnets, 10X campus behind it - $LYC (Lynas) : the non-Chinese oxides that already ship Energy Fuels + VAC: magnet craft bolted onto a US feed - Noveon: already shipping sintered magnets from Texas - Japan (Proterial, Shin-Etsu, TDK): the houses that never forgot how aibottlenecks.app/app/ai/wat… Educational, not investment advice
7
3
50
5,367
almost every high performance motor on earth turns on a sintered rare earth magnet: airpods Phone haptics EV drive units robot wrists missile fins cmpressor fans in a hall full of GPUs the bottleneck in three numbers: - China mines about 60% of the magnet rare earths - it refines about 91% - it makes about 94% of the sintered permanent magnets having deposits might be a comforting idea, it's still just a small part of the strategy. Ore is not the product. The product is basically a brick that keeps its field when the motor gets hot. Building that brick means separating periodic table twins, alloying them, powdering, sintering, machining, magnetising. China built that middle when the west treated it like a dirty job nobody wanted at home. so who owns separation, and a sinter line a customer will actually trust? A few names sit on those steps: - $MP : Fort Worth metal + magnets, 10X campus behind it - $LYC (Lynas) : the non-Chinese oxides that already ship Energy Fuels + VAC: magnet craft bolted onto a US feed - Noveon: already shipping sintered magnets from Texas - Japan (Proterial, Shin-Etsu, TDK): the houses that never forgot how aibottlenecks.app/app/ai/wat… Educational, not investment advice
7
3
50
5,367
the best model to come out this week is one that cannot write meet Jev they had a nice Wikiracing demo in the release paper. it's a game where you start on one Wikipedia page and try to reach another by only clicking links that already exist on the page. each hop: hundreds to thousands of Wikipedia links. pick one. do not invent a URL. repeat. - LLMs write a link and sometimes hallucinate a dead end (the error compounds) - Jev takes the current state, looks at the options you already defined, and returns a typed choice with probabilities. and the results speak for themselves zoom out and you can feel why this might be a paradigm shift in how we use models. A lot of software needs a reliable fork: which tool, which API, which row, which next action. If that fork gets cheap, typed, and fast, you do not run one careful agent, you spray decisions into the hot path. same shape as Wikiracing, applied to tool routing and guardrails
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
2
4
2
39
7,864
swagbench for the win hit me up with @nebiusai swag ideas in the comments, let’s see your creativity 👇
hardest eval in AI is making swag people like we need swagbench
10
1
29
7,617
okay i'll start, these cool concepts shared by @unrealromangaev
1
10
1,155
okay enough stonk talk, time for GPU talk $NBIS we just ran our widest MLPerf Inference round yet and the full rack Nebius system with NVIDIA GB300 NVL72 took first on DeepSeek R1 in both server and offline (about 603k and 690k tokens a second). simple version of what happened: 1. we got the new chips (including Vera Rubin, the latest NVIDIA iron, and we were one of only two labs to show it in this round) 2. we racked them (full GB300 NVL72, plus the smaller boxes people actually buy). 3. we put them to the test under MLPerf, the public inference benchmark where somebody else holds the stopwatch result: the full rack took #1 on DeepSeek R1 for live and batch traffic, and pushed gpt-oss 120B past a million tokens a second. When we grew from 8 GPUs to 72 (almost 9x!), speed scaled almost 9x too. One fast GPU is easy, a rack that stays fast when you multiply it is the product. another day another milestone
Five first-place results in MLPerf® Inference v6.1. 603,023 tokens/sec on DeepSeek R1 at 72 GPUs and preview-category results on the Nebius @nvidia Vera Rubin NVL72. We were one of only two submitters with results on that hardware. Full results: nebius.com/blog/posts/mlperf… #MLPerf #MLCommons #Inference
8
13
2
251
19,245
Nebius just joined @joinstationf F/ai program as a partner for the next cohort Joining alongside great names like Eleven Labs, Openrouter, Github, Hubspot, and Rippling. Station F is the biggest startup campus in Paris 🇫🇷 F/ai is the AI only track inside it: reco only, no open application, built to push early technical teams toward real revenue. the room already had OpenAI, Anthropic, Mistral, Google, Meta, Microsoft, AWS, plus Tier-1 capital like Sequoia and General Catalyst, so a european founder is not cold emailing their way into each door one by one. Cohort one ran that setup, cohort two is where @nebiusai walks in 💪 what is happening is pretty concrete. we already run GPU capacity in Paris. F/ai puts that capacity in the founder layer of the french and wider european AI scene, next to the model and capital partners who were already there. The program is also opening beyond pure Paris residency to hybrid teams across european hubs, and adding a week in SF with partner HQs. it's important bc Europe keeps producing strong research and still loses time (and sometimes companies) on the dull middle of the stack, quotas, infra help, distribution. A program like this with models, compute, tools, and capital in one coalition is trying to compress that middle. If you are building AI native in europe and the bottleneck is compute plus access, this is one of the denser rooms on the continent right now. stationf.co/news/f-ai
9
4
1
59
5,437
dylan ツ retweeted
Meet @nebius. They build cloud infrastructure for AI, including the GPU compute needed to train and run demanding AI workloads. We're putting that infrastructure in the hands of hackers for 48 hours. Barcelona brings the builders. Nebius brings the compute. Sept 19–20.
1
3
18
1,723
Two years ago Bernardo Kastrup was doing nights in an attic. Yesterday he announced a chip that consumes 100x energy maybe the next 🇪🇺 success story? dude is kind of a legend, ex-ASML, philosophy books, designed a chip aimed at neural nets that, on paper, burns a fraction of what a GPU burns because it stops dragging data back and forth across the package. This week his company raised €200m+, with Samsung as a co lead Peter Wennink (ex ASML CEO) is chair. they ship the physical system in 2028 (and license the silicon if someone else wants to build on it) Everyone agrees inference is becoming the meter. The technical problem is not mysterious, a lot of the watts never do the actual work. They pay for the trip between memory and compute. If you collapse that trip, you need fewer racks for the same token work, and basically that's the whole company. the Samsung link is quite interesting. Earlier on they were already in the manufacturing conversation, and now they are in the round. They already sell the memory, if compute and memory get designed together, that is their fight. Whether any of this works is still to see rn: no public rack, no measured watt/dollar, and until those exist, the performance claims are just claims but if that works, two things change. Cost per token stops being a GPU list price problem, and memory vendors stop being a side quest
4
3
28
4,966
if you could allocate all the compute in the world to one thing, what would it be?
37
1
43
8,400
fascinating approach periodic labs is doing something pretty clean: run materials experiments 24/7 in menlo park, use that data to train a specialist model (1T params), then put the model back into the lab to analyze xray diffraction results and help pick the next experiment. they're not trying to win a general chat leaderboard, they're trying to close a discovery loop: 1. hypothesize what material to make 2. figure out how to synthesize it 3. measure what you actually got (neon sits here hard right now) 4. feed that back into the next hypothesis They claim it beats frontier models on their xrd eval after midtrain + rl on proprietary lab data, with only 1,300 h200s for the final run (ofc treat those numbers as their eval) physical science makes this different from coding agents. You can't just spawn a thousand more experiments overnight. gear, power, and days long runs are the bottleneck. So while the wet lab is slow, they keep the gpus busy learning from data they already have. i guess the point is: the moat is the whole stack at once. the data, the model, the experiment, the lessons. You, pull one out, the rest get dumber. @LiamFedus
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
2
1
21
4,382
Elon's first Terafab milestone is not replacng $TSM i guess it's smaller and more honest: an Austin R&D line that produces something good by the end of 2027 they learn at the giga Texas research fab they multiply at the Grimes County plant 1. you usually find out whether the process works at the pilot line. 2. Then working device gives you something to measure. 3. stabilize and scale with repeated runs Only then does more equipment multiply a process you actually understand, and starts compouding to be "useful", a chip still has to clear more than one of these: - logic that works - memory that works with it - packaging that doesn’t kill either That’s why the one roof pitch matters. If logic, memory, package, and test stay split across sites, you wait on wafers and lose the loop. If they sit together, you can make a chip, test it, revise the mask, and go again without shipping the learning away. A project can be making real progress and still be nowhere near commercial volume! having a successful experimental chip answers the first question. Yield, repeatability, and the economics of turning the dial up stay open. Until you can make the useful thing on purpose (twice), you’re funding research, not inventory
* THE KOREA ECONOMIC DAILY: TERAFAB LARGELY COMPLETED PLACING ORDERS FOR FRONT-END SEMICONDUCTOR EQUIPMENT WITH MAJOR GLOBAL SUPPLIERS, INCLUDING ASML, APPLIED MATERIALS, LAM RESEARCH, AND TOKYO ELECTRON, IN Q2 THIS YEAR. SINCE THE START OF THE SECOND HALF, IT HAS BEEN PLACING ORDERS IN STAGES FOR BACK-END EQUIPMENT, INCLUDING PACKAGING EQUIPMENT. * THE KOREA ECONOMIC DAILY: TERAFAB AIMS TO COMPLETE ITS PILOT LINE BY THE END OF 2027 AND BEGIN OPERATIONS IN 2028. * THE KOREA ECONOMIC DAILY: SOUTH KOREA’S HANMI SEMICONDUCTOR IS SUPPLYING TERAFAB WITH PACKAGING EQUIPMENT FOR SYSTEM SEMICONDUCTORS.
1
1
24
6,183
the timeline went from e/acc to decel real fast
3
1
28
5,422
when a human watches a movie, he learns as it plays. A scene happens, you update, next scene, you update again. you don’t finish the film, hit rewind, and rewatch every frame on the way back. -> that’s how your brain learns, but that’s not how we train AI usually we: 1. we run the model forward 2. then we rewind the whole thing and tell every layer what it did wrong That rewind is expensive It’s why training lives in data centers, not in the device in your hand. rn the industry runs on one trick: rewind the error through the whole model. This paper might have cracked that. A 1k layer network learned from local chatter with no rewind
A training algorithm is also a bet on what computers should look like. That is how I read Sakana’s PC-ALM. The possibility that interests me is making learning practical on machines where training is currently too expensive. Sakana reports training 1,000-layer residual networks on MNIST, staying within roughly two percentage points of backpropagation. This tests whether useful learning signals can travel through a very deep network using local interactions. Sakana’s report The mechanism is elegant. In ordinary predictive coding, neighboring layers repeatedly adjust to reduce their prediction errors. In deep, narrow networks, the supervision signal can become too weak to guide early layers. PC-ALM gives each layer an accumulator that remembers its local errors during this process. That accumulated error feeds back into subsequent adjustments. Under the paper’s stability conditions, linear networks converge to the same gradients backprop would compute. Paper This separates two things that are easy to conflate: the information needed to improve a network, and the procedure used to calculate it. A familiar learning signal can emerge from a different physical process. That matters because hardware has preferences. Memory access, communication, precision and coordination all have costs. An algorithm that performs more arithmetic could still be useful if it substantially reduces something more expensive. There is already a concrete precedent. In 2022, researchers demonstrated a physical resistor network that learned through local circuit updates, without a central processor calculating those updates. These were small experiments, but the learning happened in the physical system itself. Experiment Imagine implementing repeated local adjustments directly in a circuit. The engineering question becomes how quickly and cheaply that circuit settles into a useful state. PC-ALM still has a substantial bill to pay. Its training procedure initializes with a forward pass, runs repeated local updates, then changes the weights. The iteration budget grows with depth, and the accumulators require extra memory. Algorithm and costs Shorter communication distances do not automatically mean lower total cost. Repeating a cheap operation enough times can make it expensive. But I also think there is a trap in demanding that every alternative first win on the hardware we have spent years optimizing for existing methods. We should evaluate the algorithm and its implementation together. That means measuring energy and elapsed time to reach useful accuracy, including communication, state storage and control overhead. Backprop deserves the strongest implementation we can give it in that comparison. My bet is that the first valuable application of this research direction will involve adaptation under a tight power budget. Think of a sensor learning to cope with changing conditions where it is deployed. There, the important outcome is how much useful adaptation the device can perform before exhausting its energy budget. That is a different target from adding more layers to an image classifier. It would require substantial further work, including learning over time without destroying earlier knowledge. Still, it gives this research a destination I find compelling. A device that can afford to keep learning can be useful in circumstances its designers did not fully anticipate. I would watch for the point where learning becomes cheap enough to leave switched on.
32
6,600