@meetaw

Engineering Lead @ Google Gemini. Views are personal.

Joined May 2009
Argon is a good model but this is some top tier chart crime 😂😂😂
WTF, Google just beat Claude at writing Gemini 4 Argon can write a whole novel Google's new model just took #1 in Text Arena with 1,525 - 20 points above Claude Opus 4.6, and #1 in Creative Writing The output limit: > 1M tokens per response ( up from 64K) > Opus 5.5 and GPT-6 Astra cap at 128K ~750K words in ONE GENERATION ) an average novel is ~90K ) Right now only vetted security teams have access
1
120
They have the opportunity to do the funniest thing...
1
108
Swaroop Ramaswamy retweeted
i really hate to say it, but… gemini who? 🏎️💨
Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code
471
259
232
6,550
1,655,757
Swaroop Ramaswamy retweeted
Thank you OpenAI for releasing Astra. Now Gemini 4 Pro is going to be cancelled and we'll have to wait for Gemini 4.5 Pro.
60
33
9
2,428
119,328
Swaroop Ramaswamy retweeted
I wonder how Gemini team must be feeling right now
93
14
6
864
116,447
Swaroop Ramaswamy retweeted
OpenAI: “We’ve solved math” Anthropic: “Our AI is so powerful it’s going to kill you” Google: “Introducing Gemini 3.9 Flash! It’s 30% faster and 15% worse than the last Gemini”
360
1,198
130
33,656
1,016,657
Swaroop Ramaswamy retweeted
many people say never bet against google after today's gemini 3.8 flash release but they'll never catch up
i really hate to say it, but… gemini who? 🏎️💨
8
2
1
25
15,639
Swaroop Ramaswamy retweeted
Personal Intelligence in Gemini is expanding to more people globally. 🌏 Google AI Ultra, Pro, and Plus subscribers around the world can access the feature starting today, with a rollout to free users coming soon. More information on where Personal Intelligence is available: goo.gle/422uhnd
Personal Intelligence is rolling out to more users for free across the Gemini app and Gemini in @GoogleChrome in the U.S. Access smarter responses uniquely relevant to you if you choose to connect your @Google apps like Search, @Gmail, @GooglePhotos, and @YouTube.🧵

76
161
15
1,895
255,527
Swaroop Ramaswamy retweeted
Replying to @Steve_Yegge
Maybe tell your buddy to do some actual work and to stop spreading absolute nonsense. This post is completely false and just pure clickbait.
290
432
187
12,844
858,529
So easy to switch!
New in Gemini: Import memory & chats to Gemini It's now easy to transfer your chats and personal information from other chatbots directly into Gemini Go to Settings > Import memory to Gemini "Importing memory is surprisingly smooth"
1
253
Excited for everyone to try this!
Introducing Personal Intelligence. It's our answer to a top request: you can now personalize @GeminiApp by connecting your Google apps with a single tap. Launching as a beta in the U.S. for Pro/Ultra members, this marks our next step toward making Gemini more personal, proactive and powerful. Check it out!
2
223
Swaroop Ramaswamy retweeted
Introducing a more personal Gemini, designed for you. Personal Intelligence draws insights from across your @Google apps to provide truly customized responses from Gemini. Learn all about it below. 🧵

264
381
116
2,973
560,259
🚀🚀🚀🍌🍌🍌
Made it to no.1 in the App Store. Congrats to the @GeminiApp team for all their hard work, and this is just the start, so much more to come!
3
373
I am hiring for research engineers (L4-L6) to work on building personalization for Gemini. If you're passionate about building personalized AI and have the right skillset, apply here --> job-boards.greenhouse.io/dee…. (Don't DM me, just apply directly)
2
20
2,500
This is probably my favorite Gemini feature that we have ever shipped and I don't even have kids. It. is. just. incredible.
Introducing Storybook in @GeminiApp today! Now you can create personalized, illustrated storybooks with read-aloud narration in 2 simple steps. It's free and available globally right now. Here's how to make yours + a famous story from our house about "Teddy" …
1
4
467
Grok will sometimes make mistakes, but führer over time
Grok will sometimes make mistakes, but fewer over time
233
Swaroop Ramaswamy retweeted
Many of you have been saying it, and it's true: @GeminiApp has momentum and it's growing! We've got tons in the pipeline and lots more to do!
90
101
21
1,368
140,456
Swaroop Ramaswamy retweeted
Very excited to share that an advanced version of Gemini Deep Think is the first to have achieved gold-medal level in the International Mathematical Olympiad! 🏆, solving five out of six problems perfectly, as verified by the IMO organizers! It’s been a wild run to lead this effort and I am grateful to everyone in the team for such an amazing achievement! Blog post in the thread and more to share soon!
Super thrilled to share that our AI has has now reached silver medalist level in Math at #imo2024 (1 point away from 🥇)! Since Jan, we now not only have a much stronger version of #AlphaGeometry, but also an entirely new system called #AlphaProof, capable of solving many more Olympiad problems. This is a large-scale project that I was fortunate to co-lead at @GoogleDeepMind! See our blog & NYT articel below! Blog: dpmd.ai/imo-silver NYT: nytimes.com/2024/07/25/scien…
79
221
71
1,891
654,882
5 and a half hours in and Carlos somehow found nitrous in the tank for the tiebreak. Ridiculous level from these two. #RolandGarros2025
233
Such a great feature!
We’ve been hearing great feedback on Gemini Live with camera and screen share, so we decided to bring it to more people ✨ Starting today and over the coming weeks, we're rolling it out to *all* @Android users with the Gemini app. Enjoy! PS If you don’t have the app yet, download it here: goo.gle/4imLD45
267