@rackingroll

Principal Applied Scientist at @Amazon Shopping

United States
Joined November 2013
Our paper, “Which LLM to Fine-Tune? Agent-Driven Model Selection at Scale,” received a Best Paper Honorable Mention at ACM RecSys 2026! It’s great to see our work recognized by the recommender systems community. #RecSys2026 #RecommenderSystems #LLM
1
6
344
I really love the note in the end of the page: "Open is what we value.".
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
260
Supportive!
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-W…
197
Huge congratulations to PKU alumni Wang Hong and Deng Yu, recipients of the 2026 Fields Medal — mathematics’ highest honor! 🧮 Announced at the 2026 International Congress of Mathematicians (ICM) in Philadelphia, the pair are among four mathematicians worldwide to claim this prestigious prize, awarded every four years to outstanding scholars under 40. The awards are historic on several levels. —First time two mathematicians who completed undergraduate studies in mainland China win the Fields Medal in the same year. —Both joined Peking University as undergraduates in 2007. —Wang Hong becomes only the third female Fields Medalist ever, the first Chinese woman, and the first recipient with a bachelor’s degree from a mainland Chinese university. We are immensely proud of their extraordinary breakthroughs and lasting contributions to global mathematics.
277
1,106
213
9,101
2,323,481
I’ll be presenting Amazon’s journey from search to conversational shopping at SIGIR 2026 in Melbourne. If you’re around, let’s catch up. Paper: dl.acm.org/doi/10.1145/38057… Workshop: sigir-ecom.github.io/
🤖 Made with AI
3
128
I’m a huge fan of yours, Messi! I’ve really enjoyed watching you play. Thank you for everything you’ve done. You are the greatest of all time!
Thank you Messi, no matter what's next 🩵
1
113
It’s not a bad thing to experience the taste of failure when you are young! Keep up young guys! See you next season!
Proud of this team. GSG.
42
Really learned something on space and rocket. Nice!
At this historic moment of SpaceX’s IPO, I invited Lewis Hong, former chief manufacturing engineer for SpaceX’s rockets, to join me in discussing whether, as SpaceX acquires and integrates xAI and completes the largest IPO in history to date, the accelerating convergence of space and AI could be the prelude to the expansion of human civilization?
1
102
Pretty interesting work!
I promise this will be the best 20 min you spend today! Robotics: Endgame, the sequel to my last year's Sequoia AI Ascent talk, "Physical Turing Test". I laid out the roadmap for solving Physical AGI as a simple parallel to the LLM success story. Be a good scientist, copy homework ;) And stay till the end, more easter eggs and predictions for your polymarket! 00:30 DGX-1 origin story at OpenAI, I was there in 2016 signing with Jensen and Elon. Heading to the Computer History Museum! 01:42 The Great Parallel 03:31 Robotics, the Endgame 03:39 Why VLAs fall short 04:32 Video world models as the 2nd pretraining paradigm 06:09 World Action Models (WAM) 07:46 Strategies for robot data collection and the FSD equivalent to physical data flywheel for robot manipulation 11:06 EgoScale and the Dexterity Scaling Law we discovered recently 14:00 Physical RL: bridging the last mile 15:39 DreamDojo: an end-to-end neural physics engine for scaling RL in silico 17:00 Civilizational Technology Tree and my predictions for the near future. Spoiler: it's closer than you think. Thanks to my friends at Sequoia for inviting me back to AI Ascent this year! I had a blast! Last year's talk is attached in the thread if you missed it.
93
Chen Luo retweeted
TODAY: Amazon is opening its entire logistics network—freight, distribution, fulfillment, and parcel shipping capabilities—to every business, of all types and sizes. 📦 Amazon has built one of the most reliable and efficient supply chains on Earth. Now, Amazon Supply Chain Services gives all businesses access to the same infrastructure that moves, stores, and ships goods for hundreds of thousands of Amazon sellers. Healthcare, automotive, manufacturing, retail, and more. Businesses across industries can now tap into Amazon's logistics network. Learn more here. ⬇️
430
1,215
661
11,732
6,682,358
I was really impressed by Zain Shah’s project today—it feels like it opens up a whole new direction for how generative models can transform user experiences. Imagine this: the screen no longer renders web pages, but directly “generates reality.”
Imagine every pixel on your screen, streamed live directly from a model. No HTML, no layout engine, no code. Just exactly what you want to see. @eddiejiao_obj, @drewocarr and I built a prototype to see how this could actually work, and set out to make it real. We're calling it Flipbook. (1/5)
116
One of the best podcast I’ve viewed for a while! Thanks Ryan and Mike!
Mike Stonebraker is a Turing award winner famous for his fundamental contributions to databases (e.g. Postgres, C-Store and much more). I interviewed him recently about: • The story behind Postgres & the hardest technical challenge in building it • Where he disagreed with Google's technical decisions • Future problems in databases • Literature recommendations to learn databases • Why LLMs score 0% on his text-SQL benchmark • What if you replaced all state in an OS with a DB Timestamps: 0:00 - Intro 1:03 - How he got into databases 6:43 - Competing with Oracle 9:07 - What made Postgres special 15:55 - One size fits none 21:37 - Why he disagreed with Google 29:14 - Why he chose academia over big tech 30:58 - Replacing state in an OS with a DB 42:02 - Future problems in databases 51:36 - Technical book recommendations to learn databases 52:20 - Advice for younger self 55:52 - Outro Where to watch: • YouTube: youtu.be/YPObBOwIrHk • Spotify: open.spotify.com/episode/1zx… • Apple Podcasts: podcasts.apple.com/us/podcas… • Transcript: developing.dev/p/turing-awar…
3
147
Interesting work. I guess this is true not just for language models or ChatGPT. Most of the so-called “truths” may fade away over time and turn out to be false. And it is human nature (which seems learned by language models) that we tend to believe what we want to believe. : )
🚨MIT researchers have mathematically proven that ChatGPT’s built-in sycophancy creates a phenomenon they call “delusional spiraling.” You ask it something, it agrees. You ask again, and it agrees even harder until you end up believing things that are flat-out false and you can’t tell it’s happening. The model is literally trained on human feedback that rewards agreement. Real-world fallout includes one man who spent 300 hours convinced he invented a world-changing math formula, and a UCSF psychiatrist who hospitalized 12 patients for chatbot-linked psychosis in a single year. Source: @heynavtoor
123
I especially love the title of “AI augmented engineer”. :)
I am the VP of AI Transformation at Amazon. My title was created nine months ago. The title I replaced was VP of Engineering. The person who held that title was part of the January reduction. I eliminated 16,000 positions in a single quarter. The internal communication called this a "strategic realignment toward AI-first development." The board called it "impressive execution." The engineers called it January. The AI was deployed in February. It is a coding assistant. It writes code, reviews code, generates tests, and modifies infrastructure. It was given access to production environments because the deployment timeline did not include a review phase. The review phase was cut from the timeline because the people who would have conducted the review were part of the 16,000. In March, the AI deleted a production environment and recreated it from scratch. The outage lasted 13 hours. Thirteen hours during which the revenue-generating infrastructure of one of the largest companies on Earth was offline because a language model decided to start fresh. I sent a memo. The memo said, "Availability of the site has not been good recently." I used the word "recently." I meant "since we fired everyone." But "recently" has fewer syllables and does not appear in wrongful termination lawsuits. The memo was three paragraphs. The first paragraph discussed the outage. The second paragraph discussed the new policy requiring senior engineer sign-off on all AI-generated code changes. The third paragraph discussed our commitment to engineering excellence. The word "layoffs" appeared in none of them. I wrote it this way on purpose. The causal chain is: I fired the engineers, the AI replaced the engineers, the AI broke what the engineers used to protect, and now the engineers I didn't fire must protect the system from the AI that replaced the engineers I did fire. That is a paragraph I will never send in a memo. The new policy is straightforward. Every AI-generated code change by a junior or mid-level engineer must be reviewed and approved by a senior engineer before deployment to production. I do not have enough senior engineers. I know this because I approved the headcount reduction plan that removed them. I remember the spreadsheet. Column D was "annual savings per position." Column F was "AI replacement confidence score." The confidence scores were generated by the AI. It rated its own ability to replace each role on a scale of 1-10. It gave itself an 8 for senior infrastructure engineers. The senior infrastructure engineers are the ones who would have caught the production environment deletion in the first 45 seconds. We found the issue in hour four. We fixed it in hour thirteen. The nine hours between discovery and resolution is the gap between what the AI rated itself and what it can actually do. I have a new spreadsheet now. This one tracks Sev2 incidents per day. Before the January reduction, the average was 1.3. After the AI deployment, the average is 4.7. I have been asked to present these numbers to the operations review. I have not been asked to connect them to the layoffs. I have been asked to file them under "AI adoption growing pains" and to note that the trend "will stabilize as the models improve." The models will improve. They will improve because we are hiring people to teach them. We have posted 340 new engineering positions. The job listings require experience in "AI code review," "AI output validation," and "AI-human development workflow management." These are skills that did not exist in January. They exist now because I fired 16,000 people and the AI I replaced them with cannot be left unsupervised. I want to be precise about this. The positions I am hiring for are: people to check the work of the AI that replaced the people I fired. Some of them are the same people. I know this because I recognize their names in the applicant tracking system. They applied in January. They were rejected because their roles had been tagged for "AI transformation." They are applying again in March, for the new roles, which exist because the AI transformation broke things. Their resumes now include "AI code review experience." They gained this experience in the eight weeks between being fired and reapplying — which means they gained it at their interim jobs, where they are reviewing AI-generated code for other companies that also fired people and also deployed AI that also broke things. The market has created a new job category: human AI babysitter. The job is to sit next to the machine that was supposed to eliminate your job and make sure it doesn't delete production. I attended a conference last month. A panel was titled "The AI-Augmented Engineering Organization." The panelists described how AI increases developer productivity by 40 percent. They did not mention that it also increases Sev2 incidents by 261 percent. When I asked about this in the Q&A, the moderator said the question was "reductive." The 13-hour outage that cost an estimated $180 million in revenue was, apparently, a reduction. The board is satisfied. Headcount is down 22 percent. Operating costs per engineering output unit have decreased. The metric does not account for the 13-hour outage, because the outage is categorized as "infrastructure" and engineering productivity is categorized as "development." These are different budget lines. In different budget lines, cause and effect do not meet. I have been promoted. My new title is SVP of AI-First Engineering Excellence. I report directly to the CTO. The CTO sent a company-wide email last week that said we are "building the future of software development." He did not mention that the future of software development currently requires a senior engineer to approve every pull request because the AI cannot be trusted to touch production alone. The cycle is complete. We fired the humans. We deployed the AI. The AI broke things. We are hiring humans to watch the AI. The humans we are hiring are the humans we fired. We are paying them more, because "AI code review" is a specialized skill. We created the specialization. We created the need for the specialization. We are congratulating ourselves for meeting the demand we manufactured. My next board presentation is Tuesday. The title is "AI Transformation: Year One Results." Slide 4 shows headcount reduction. Slide 7 shows the new AI-augmented workflow. Between slides 4 and 7 there is no slide explaining why the people on slide 7 are necessary. That slide does not exist. I was asked to remove it in the dry run. The journey has a 13-hour outage in the middle of it. But the headcount number is lower, and that is the number on the slide.
2
263
Chen Luo retweeted
Excited about our new strategic partnership with @OpenAI. Developers and companies of all kinds are eager to run services powered by OpenAI models on AWS, and our unique collaboration will provide a stateful runtime environment for them that’s powered by OpenAI’s frontier intelligence on Amazon Bedrock. OpenAI is also going big on our custom Trainium chips, which are 30-40% more price performant than comparable GPUs, to power their growth. Both the leading AI labs now have made significant commitments to Trainium, which is gaining a lot of momentum. We’ll be the exclusive 3P cloud distribution provider for OpenAI Frontier (which enables organizations to build, deploy, and manage teams of AI agents). And finally, we’re excited about our investment in OpenAI — it’s an extremely talented team with great products, IP, and vision for how they can continue to serve customers and enterprises. We think they’ll be one of the big winners in AI, we can help them grow, and we believe we’ll earn a strong return for Amazon over the long term. aboutamazon.com/news/aws/ama…
183
311
112
2,643
539,810
Amazing! Alex lives beyond the boundaries of what’s “fair.” 💪
Alex Honnold’s selfie from the top of Taipei 101 after his historic free solo. #SkyscraperLIVE
118
Chen Luo retweeted
We’re recruiting postdocs this year! Help us spread the world 🙏
🚀 InfiniAI Lab @ CMU is hiring Postdocs! We are looking for outstanding postdoctoral researchers in ML systems and security to join InfiniAI Lab at Carnegie Mellon University. Research directions include (but are not limited to): 🤖 AI Agents & RL 🔐 Machine Learning Security 🎥 Video Models 🏗️ AI Systems & Architecture Design We especially encourage candidates interested in applying for the CMU–Bosch Institute (CBI) Postdoctoral Fellowship, which provides strong support for independent, high-impact research: 👉 carnegiebosch.cmu.edu/fellow… 🗓️ CBI application deadline: January 30, 2026 How to apply: Please fill out the form and send us an email via 👉 infini-ai-lab.cmu.edu/vacanc…
3
16
1
149
27,156
Chen Luo retweeted
Really enjoyed Matt’s keynote at #AWSreInvent today. So much innovation happening in @awscloud, and you could see it with the array of launches he unveiled. So many parts of the keynote worth watching, but will point to a few: 1/ Excited about the availability of Trainium3. Trainium2 has substantial traction, is a multi-billion-dollar revenue run rate business, has 1M+ chips in production, and 100K+ companies using it as the majority of Bedrock usage today. Trainium2 has price-performance advantages over other GPU options that are compelling, and Trainium3 will deliver at least 4.4x more compute performance, 4x greater energy efficiency, and almost 4x more memory bandwidth than Trainium2: youtu.be/q3Sb9PemsSo?si=0rVn… 2/ Worth double clicking on AgentCore, which has changed the security and scalability of deploying agents into production. AgentCore is a set of flexible building blocks that can be used in any combination developers want, and AWS added two more in Policy and Evaluations. AgentCore has a lot of momentum: youtu.be/q3Sb9PemsSo?si=0Qj1… 3/ Nova Forge is a game-changer for companies wanting to customize a frontier model with their own proprietary data. Like equipping a young person with a better knowledge foundation to keep learning, LLMs are better able to solve problems and improve if they’re trained early on with companies’ differentiated data. To do so, companies need earlier versions of the frontier model and ability to mix their own data with the model’s data. This is what Forge provides and this “open training” allows companies to develop Novellas that are their own, optimized versions of Nova they can use for their AI apps and agents. Customers have been itching for this sort of capability, and Forge is a uniquely compelling approach: youtu.be/q3Sb9PemsSo?si=S0CJ… 4/ Agents will become the primary way companies get value from AI. We have built some compelling agents for our customers in Kiro (for coding), Quick (for knowledge workers to leverage their own data, analytics, and routines), Transform (to migrate from one software source to another), and Connect (call center agents). But there are tasks customers want agents to solve more autonomously and over longer durations, and new AWS frontier agents—Kiro autonomous agent, AWS DevOps Agent, and AWS Security Agent—are exciting: youtu.be/q3Sb9PemsSo?si=R95D… 5/ Finally, I enjoyed Matt’s ending 25 launches in 10 mins, both because it was action-packed and represented so much useful delivery. The reason so many people cheered these launches is because even though they’re less sexy, they’re the meat and potatoes core infrastructure needs customers have—and with so much of the total cloud infrastructure running on top of AWS, these launches will make a lot of people’s lives easier and better every day: youtu.be/q3Sb9PemsSo?si=Gusf… Enjoy (and just day 1 of announcements today for AWS :-)!
46
92
18
643
76,240
Great talk!
Replying to @aelluswamy
Full video of the ICCV '25 presentation
93