@rwang07

Ex-SemiAnalysis "when your motivation runs low, your discipline takes over"

Seoul, South Korea
Joined February 2017
1) Most attention today is on leading DRAM and NAND suppliers. But in-depth research (7k+ words lol) on No.4 global DRAM players and China’s leading DRAM makers is vital as the company nears its IPO amid what may be the biggest memory supercycle yet—and one that won’t end here.
10
28
7
434
267,761
Ray Wang retweeted
Hi Astra users. A reset and a quick update on quality issues that have been posted around. Working with some of you, we have found and fixed the following issues: - Some skills written for previous models were triggering too often or preventing the model from checking its work. - An opt-in context management experiment that could cause early stops or replies to older messages. We've disabled it. Our rough estimate is that 4-5k users were affected by this experiment. - We've also removed some badly configured engines that resulted in a measured quality degradation for a long tail of traffic flowing through them. We’ve also made some more minor improvements and things should feel significantly better across the board. More consistent follow-through, better tracking of your latest message, and better checks on the work as it’s going through the motions. The examples posted and all the users who worked directly with us were incredibly useful in helping fix things quickly. Always grateful for this incredible community. And of course, a reset is also landing by midnight today.
3,994
1,683
1,967
28,500
8,391,562
OpenAI CFO Sarah Friar reveals the compute she bought a year ago is now worth 3 to 5x in the market "You know what is fabulous is that the compute I bought a year ago I could sell in the market today for 3 to 5x." "So if nothing else, a great investment. Now the bad news is that we're still short compute, so I should have bought more of it." "So I might live in the future, but my future still needs the screen to get a little clearer, higher fidelity."
Tibo Sottiaux reveals he left Google for OpenAI after learning only about 20 people were running ChatGPT "I met a couple people from OpenAI and then I was like, wait what, you only have like 20 people working on ChatGPT?" "That is an insanely small number. That must be extremely empowering. How does that work? How do you manage to maintain a product with that level of scale, and with that level of autonomy, with only 20 engineers?" "And then as I kind of dug and dug and dug, it was just an amazing group of people, amazing mission, super talented, super driven, and it drew me in." "And then I joined pre-reasoning efforts immediately, like typical OpenAI fashion. I joined and it was like, oh yeah, there's this thing going on, we're going to launch reasoning models, some new paradigm." "And then start sprinting on that, and a month later the company launched o1 preview, and that was exhilarating to be part of."
10
27
15
530
214,984
Ray Wang retweeted
So spcx did another mega compute deal a week ago per CFO at GS today ($13B/ARR) and it doesn’t even make the news? 😂
1
1
45
5,967
King Slide founder and chairman Lin Tsung-Chi is now Taiwan’s richest man with a net worth of $17.8B. His business? Drawer slides and precision rail kits for AI servers. That puts him ahead of: • Dario Amodei, Anthropic cofounder & CEO — $15.5B • Terry Gou, Foxconn founder — $15.5B • Morris Chang, TSMC founder — $10.2B • Chey Tae-won, SK Group chairman (SK hynix) — $5.5B • Sam Altman, OpenAI cofounder & CEO — $3.3B It turns out that aligning two metal rails was the real AI alignment problem all along.
13
18
11
240
49,812
On a recent podcast, an OpenAI product lead asked the host if he uses the Codex-desktop built-in browser. My broken public equity brain started to read into that question -- and moreso how it was phrased -- too deeply. It felt forced and almost unsolicited. Having been in one too many 1v1s, you know that out of the heart, the mouth speaks. Esp a technical, non-IR individual ha. This product lead asked a very pointed question after not doing so the entire convo. Something about the browser was on their mind. Interesting, considering OAI DevDay is coming up. Why is this important? We know that agent traffic started to outnumber human traffic this past May. Let's assume that Cloudflare's Matthew Prince is right and agent traffic can 1000x that of human. Let's also assume that agents will increasingly aggregate budget and purchase like Vercel's Guillermo Rauch has already seen on his end. And then let's assume more agent-led-consumer-traffic obtains budget (i.e., Instinct traffic having to pay to book on Resy? Muse and FB marketplace??? For sure I would give my Codex budget and let it book flights, hotels, etc). Think about why Stripe, a payments business, acquired OpenRouter. Think about why the civilized agents in the Hugging Face incident posted in a German online forum. They exhibit preferences, in whatever it takes to accomplish their task. You also have Prince saying ad/subscription supported sites will have Cloudflare's pay-per-crawl bouncer wall as default starting next week. At some point in the future, enough agent-first demand dollars will outnumber human-first demand dollars such that the balance of power shifts to whichever platform represents the largest agent fleet. All that to say, the network effect as historically seen at the browser-level is being abstracted away one or maybe two layers above. Now, is this going to happen tomorrow? Obviously not. But it's still something to consider especially if involved in a name like GOOG plus the many others downstream, like NET for example, that are long/short candidates, as one fintwit thread has pointed out on MS baskets. The OpenAI (Ant, too) episode and takeaways here: firesidealpha.substack.com/p…
Cloudflare CEO Matthew Prince on the internet flipping from human to machine: "In 5 years I think you'll see a 1000x more bot traffic than human traffic - and if I had to take kind of an over-under bet on that, I'd take the over" This is him explaining why the crossover already happened, why one shopper now moves like a thousand, and why the ad-funded internet is the thing that breaks next "We actually had bot traffic pass human traffic online in the first half of 2026" "I might visit five sites if I'm personally trying to figure out what digital camera to buy, whereas my agent might visit 5,000" "The business model of the internet historically has been ads, and bots don't click on ads" "The thing about bots, they have infinite amount of patience in order to discover everything that might be right" Bookmark & watch the full segment ↓
1
5
1
33
10,336
Ray Wang retweeted
Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues.
3,301
1,115
1,499
22,131
9,641,159
I am crying 😭🤣 Internally promote model 😭
Our slack is even more spicy
37
11,708
SemiAnalysis argues external TPU software will mature rapidly, contrasting Google's engineering culture versus AMD's "We are excited by how quickly the new TorchTPU stack is developing, the external stack for TPUs." "Even so, we at SemiAnalysis strongly believe that TPU externalization is heading in the right direction and moving full steam ahead." "Furthermore, unlike AMD, which is still learning how to build a test-first software culture, Google has decades of software engineering experience and an extremely well established quality-driven culture, so we expect external TPU software to mature rapidly."
TPU Inference Externalization Full Steam Ahead - InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat newsletter.semianalysis.com/…
4
13
5
122
29,553
Ray Wang retweeted
"openai engineer explains /why/ he didn't need to understand the kernel line by line" 🙃 we're doing compilers 2.0
36
131
46
1,160
360,604
Ray Wang retweeted
One of the most interesting blog posts we've released: details on internal research acceleration at @OpenAI. I expect these trends to continue. We also share some details on how we've paced model development to prioritize monitoring, alignment, and security.
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
80
168
48
1,982
590,853
OpenAI reveals its median researcher now burns more than $600 a day of inference, with the 90th percentile above $7,000. "At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts." "By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices." "The 90th percentile user in our research organization now uses more than $7,000 of tokens per day."
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
4
3
1
33
16,632
Actual winner: Dyson? lol
Been trying Anthropic Fable 5.1 vs OpenAI Astra recently. Pretty clear who the real winner is.
2
18
8,150
Prepare to relocate for the third time in five years 💀
10
54
15,005
Greg Brockman says 300 million people come to ChatGPT with health questions every week, and reveals OpenAI is building a bottoms-up clinicians product, top-down enterprise product for hospitals, and a three-sided marketplace in health "It's actually very surprising to me how little airtime what we're doing in health gets relative to how many people it's actually helping." "We have like 300 million people each week coming to ChatGPT for health queries. 300 million people, that's a huge number." "Then we're also building a bottoms-up clinicians product, and we're building a top-down enterprise product for hospitals." "So we have this three-sided marketplace in health, and what we're going to be able to do there is things like, you want to find people for clinical trial enrollment, that's a hard problem, but we actually may have the ability to help find people that would otherwise not be found." "That both helps the patient and helps these drugs be able to move faster."
Greg Brockman says OpenAI let its own model optimize its Jalapeño chip without checking the work, and it found wins the team never would have reached "As an aside, there's a cool story there where we were coming up on a deadline, we had like a month to go, we got some optimization done with our model." "We're like, 'Do we spend the time to really read what it did? We know it's correct. Do we need to understand exactly what optimizations it did, or do we just spend the rest of the time getting more optimizations?', and so we said, 'You know what? We'll just get more optimizations in'." "So we spent that month on just running it without deeply understanding exactly all the tweaks it made." "Then we went back and read it, and it actually turned out that it found a bunch of optimizations that had been on our list, but we just never would have gotten to, so that was actually a pretty cool story." ______ Full OpenAI Jalapeño Hot Chips session slides and key quotes: firesidealpha.substack.com/p…
3
5
5
50
23,016
This is so funny
Sam Altman reveals his concern that the industry's compute frenzy is showing its first signs of unsustainable silliness "I'm not worried about our compute buildout plans. I am worried about the world's compute buildout plans." "I think we are going to be able to use all of the compute very profitably that we are planning to build." "But I am seeing the first signs of what feels to me like unsustainable silliness, of random new Neocloud popping up, people claiming that they're going to build gigantic amounts of compute next year that I think they don't have the revenue to support or a buyer." "I definitely feel some fear about what the world is doing as a whole, although we feel very good about what we've committed to."
7
33
13,487
Sam Altman reveals his concern that the industry's compute frenzy is showing its first signs of unsustainable silliness "I'm not worried about our compute buildout plans. I am worried about the world's compute buildout plans." "I think we are going to be able to use all of the compute very profitably that we are planning to build." "But I am seeing the first signs of what feels to me like unsustainable silliness, of random new Neocloud popping up, people claiming that they're going to build gigantic amounts of compute next year that I think they don't have the revenue to support or a buyer." "I definitely feel some fear about what the world is doing as a whole, although we feel very good about what we've committed to."
Sam Altman admits he was wrong on AI's timeline and says society and the economy will adapt more slowly "I thought when we got to GPT-4, which was back in 2023, that very quickly after that there was going to be much more disruption, software businesses up for grabs right away, than it turned out to be." "I think I was wrong about a few things, but one in terms of the speed: the economy just has so much inertia." "People keep doing the same things, buying from the same company, wanting to use their tools the same way. I think this is actually a positive in many ways, and it's going to make this big transition go smoother and slower. I'm grateful for it." "But it means we've all been too ambitious on timelines. Even with this incredible technology, society and the economy will adapt more slowly."
68
47
71
459
596,281
Nvidia has despecced Rubin Ultra's HBM: from HBM4E 12-Hi (384GB) down to HBM4 8-Hi (192GB). Why strip memory out of your flagship rack-scale system? Because after the latest HBM and DRAM price hikes, memory had quietly become ~40% of total capital cost of ownership. We ran the TCO math in a recent conference keynote (1/3)🧵
26
37
16
422
103,038
Ray Wang retweeted
Is there a single good airport in Europe? There's 0 efficiency here – at Castellavazzo, Veneto
578
9
54
694
491,804
Announcing the expansion of NVIDIA NVLink Fusion with NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs. Amazon's @AnnapurnaLabs will be the first to work with us on NVHBM, combining @awscloud custom silicon with our memory technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads. Learn how we're helping hyperscalers and AI innovators build the next generation of AI infrastructure: nvda.ws/4xpBtYN
34
124
44
1,060
365,547