🚨Prediction: all decode silicon is going to be sold out for a decade at least.. follow the 🧵....
@nvidia reportedly plans to invest in @dMatrix_AI , an inference chip company (and our partner) that on paper is a competitor. A month ago, d-Matrix announced its Raptor XPUs will plug directly into Nvidia MGX racks over NVLink Fusion.
Why this matters:
1/ Companies investors thing are competitors to Nvidia become co-decode partners. GPUs handle prefill and verification. Specialized silicon like Raptor handles decode and draft models. The whole workload stays inside an Nvidia AI factory instead of going to rival GPUs or hyperscaler chips.
2/ Jensen wants a bigger pie. We’ve heard from inside Nvidia that the only thing Jensen really cares about is growing the market. That’s why Nvidia backs open source, robotics and world models, and spends thousands of engineering hours expanding the AI TAM. d-Matrix is the same playbook applied to inference.
3/ Fast decode is the unlock. Decode is memory bound, and dedicated decode silicon (or a tonne more memory in GPU racks) is the only practical path to fast inference. Fast inference creates new workloads the way broadband did after dial-up. Nobody built Netflix on a 56k modem.
4/ The real bet is memory. D-matrix's new solution - Raptor - bonds custom 3D DRAM face to face with a 4nm logic die, with no PHY in between. d-Matrix measured 0.37 pJ/bit on early silicon and is targeting 100+ TB/s per card. Its HBM4 reference point: ~18 TB/s at 2 to 3 pJ/bit. Memory bandwidth is the bottleneck for decode. By backing d-Matrix, Jensen is helping make sure memory innovation reaches the market, and that it reaches it inside NVLink.
Racks land Q4 2027. The heterogeneous AI factory is coming, and Nvidia intends to be the platform it runs on.
At @general_compute , this is the future we’re building for.
@JensenHuang is the GOAT.
Woah, @sidsheth at @dMatrix_AI making the big moves
You know the future of compute is heterogeneous when @JensenHuang
starts putting down some chips
Now we know why they integrated into nvidia's nvlink a few weeks back
P.s you could have mentioned this at dinner last night buddy 😄
Over 1,000 person waitlist for our #sftechweek event today
Top of @a16z recommended events this week
If you are a vc, and follow me, don't sleep on this announcement.
This isn't another memory heavy approach to decoding fast ai.
This isn't another rack scale solution to lower TCO.
This is a new compute paradigm.
It will be hard to get right.
But if it works, it fundamentally changes the entire compute landscape and history of intelligence.
@extropic @beffjezos
Thermodynamic Computing: From One to One Billion
0:04 - Intro
1:22 - Recap
2:04 - Torx
3:37 - Thermalizers
4:42 - Hardware
6:45 - API
7:25 - Conclusion
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Yes.
Cyber wars results will be a function of model intelligence + latency.
We already teamed up with @wafer_ai to provide the worlds fastest k3 in a bitcoin hack defence clean up.
This is where the market is heading in 12-24 months.
Finn Puklowski retweeted
are there neoclouds offering bare-metal SRAM accelerator access? (not smth like groq api)
I think @gimletlabs-esque multi silicon inference makes a ton of sense. and autonomous kernel support across the chips. would be cool to try at home.
awesome first day all together in sf for our first offsite with the @general_compute team.
Looking forward to quality meetups with friends and partners this week.
SemiAnalysis numbers imply OpenAI is selling Cerebras-powered Ultrafast inference at ~$200M per megawatt per year.
People are re-running the maths hoping for a typo. We're not surprised.
A megawatt is a megawatt. What it earns depends on the silicon inside it.
Here's how we think about it:
Near term, revenue per megawatt wins. Power is the scarcest input in AI, and whoever monetizes each MW best gets the next MW.
Long term, gross margin per megawatt wins. Prices compress, competitors arrive, and only the operators with real cost and pricing advantages keep the spread.
Cerebras is built for both.
Speed is the product. A frontier model at 750 tokens per second is something GPU clouds haven't publicly matched, and customers like Jane Street pay a premium for it. Premium pricing on a contracted, scarce resource is how you protect margin when everyone else is racing to the bottom.
That's why we're extremely bullish on Cerebras, and why we're buying as many megawatts as we can.
Premium inference unlocks the most delightful user experience in AI. No more waiting, waiting, waiting, staring at little thinking animations while the model works in the background. Answers that keep up with your thoughts.
Our customers are even more excited than we are.
That's General Compute.
love this analogy from @evanjconrad
pooling demand is how electricity got cheap. and cheap electricity didn't mean we used less of it, we plugged everything in
same with compute. the bottleneck is real, especially for really performant compute. but as the speed of inference goes up and the cost goes down, demand goes up with it
IMO Jevon's law will hold true
expect it to continue...
nitter.cf/tbpn/status/2104709092…
SF Compute CEO @evanjconrad compares today’s compute market to electricity before the power grid.
He explains that before utilities pooled demand, factories had to build enough generation for their own peak usage, leaving expensive capacity idle the rest of the time.
“Every factory ran its own generator. And they had to size that generator for peak capacity. Then during idle hours they would just sort of burn that money.”
“What happened in the power markets is that utilities came in and they pooled demand between all of these different factories.”
“By doing that, you made the economics of all the factories better.”
“This is how we scaled up electrification in the country.”
“This is what people are trying to do with compute markets.”
when i was younger i used to be afraid a competitor would copy my business
and actually with my last company, one of them tried
the tldr is that ideas are worth jack,
the same applies to asics and shows why nvidia's moat has lasted for so long
it's nailing the operations, practicalities and distribution that make something work
that is genuinely hard and most people find that out pretty late
The pace of change in the silicon world is crazy rn
A year ago 3 of the 15 inference asic programs we looked at had shipped silicon in volume.
Now that number's 9. Each of them have their special sauce or angle.
Some compete with Nvidia - some compliment like the memory focussed folks
Mini-Cambrian explosion happening and i'm quite sure most of them will be roaring successes
Shipping in volume with failure testing, rack scale solutions is where our lenders, not just a VC, can underwrite the growth at scale
This year is going to be huge. Next year even huge-er
We first met the team at @Cerebras before we'd even named our company
Almost a year later, we're making a huge bet on wafer-scale
Excited to see where this partnership heads 🚀
nitter.cf/general_compute/status…
🚨 BREAKING: Super excited to announce we're deploying the world's fastest inference with @Cerebras.
Talk to any developer and they're excited to build with 20x faster AI... the problem is there's almost no compute available.
We're here to solve that.
Using Nvidia GPUs for prefill - it's now more affordable than ever too.
Thank you to the whole Cerebras team and excited to grow this partnership.
First tokens live Q127 🚀
There's a new version of this post
Interesting experience in running our latest raise
There's VCs that get compute, and then there's VCs that are still in the b2b saas model
The former conversations are a pleasure
The latter are tbh, painful
I find it hard to fathom how hard it must have been to do deep tech / hardware raising only 4 years ago
Hats off to you founders that built the companies like @SambaNovaAI, @cerebras etc before it was cool
Is it then possible that raising rates is inflationary for the compute stack as cost of capital flows down to higher compute prices per hr - at least in our little domain it would mean higher customer prices
The presumption that the Fed raising short-term rates reduces inflation is predicated on the belief that higher rates reduce demand and investment.
But what if higher rates don’t reduce demand and investment because the demand for intelligence and energy is unaffected by higher rates because winning the race for super intelligence has a near infinite ROI and the demand for compute will remain incalculable.
Why won’t higher rates at this unique moment in history therefore lead to more inflation as interest costs are embedded in everything?
And the problem is compounded as the more the Fed raises rates, the more inflation we will have and the more the Fed will need to raise rates further and so on.
But what if the old models don’t apply to the current paradigm and the Fed is wrong?
I think the Fed might have just made a mistake. Am I right or am I wrong?
couldn't agree more,
that means they must match model intelligence as well as inference speed
cyberdefense loses on speed, now that intelligence is a commodity
maybe I should really start this programme to give free compute to cyber defence teams
nitter.cf/ilyasut/status/2094881…
came across this and this is the trend we're seeing across our customers..
as llms get better, customers usage rates increase -> higher NPS -> more customer referrals
it's a usage flywheel
nitter.cf/gsivulka/status/209515…
Weeks after launching Max, we’re growing faster than at any point in our history.
In the last quarter:
• DAUs have 3Xed…
• Total LLM usage has 5Xed, across both frontier, and our own post trained models…
• 40%+ of monthly active users use Hebbia on an average day, and 70%+ use it in a given week, with many of our customers still onboarding to the latest product.
The business has never been stronger:
• We beat 7 competitors to win the largest RFP in investment banking, and closed 10 banks shortly thereafter expanding our TAM well beyond our core investor business.
• We hired 100 people since January, doubling our headcount across London, SF, and NYC.
• We launched a new SKU that closed several of the biggest contracts in company history within *weeks* of first customer conversations.
• The majority of our Series B raise remains untouched. We remain remarkably capital efficient.
There’s a lot of noise in the market. This is a very dynamic market that changes rapidly, and we’ve never been more bullish on our market adoption, product leadership and ability to execute moving forward.
And of course, we’re hiring. careers.hebbia.ai
Come build the future of finance.
love this.
first computer and it boots risc-v, so @amasad has him on alternative silicon before he can spell it 😄
nitter.cf/amasad/status/21018522…