@meetp_ai

🇮🇳 | Computer Vision Engineer | AI Research

Joined June 2015
I write about frontier AI models, coding agents, benchmarks, and the product/policy decisions shaping AI access. Mostly: GPT/Claude/Gemini, agentic coding, evals, and real-world AI tooling.
288
Starfall Improved Graphics fidelity and gameplay.
Astra rebuilt it in 3D with barrel rolls, boosts, a boss battle and sound effects. Then I asked it to prepare the new demo video. Here’s the result 🔊
26
Astra rebuilt it in 3D with barrel rolls, boosts, a boss battle and sound effects. Then I asked it to prepare the new demo video. Here’s the result 🔊
Made Starfall type game in Single Prompt with Astra (high) Demo Gameplay video also prepared in the same prompt.
1
235
Made Starfall type game in Single Prompt with Astra (high) Demo Gameplay video also prepared in the same prompt.
Replying to @reach_vb
Made it cooler.
1
1
306
Now we know OpenAI’s side of the full story. Not sure who’s right or wrong here, but Seb does look like the more transparent party here.
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
55
Heard a rumour about a math problem. 10,000 concurrent agents and 88 hours later. Navier-Stokes Millennium Prize Problem! Solved!!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
19
AI Benchmarks are playing catch up with the pace of capability jumps from both the frontier labs. And they’re currently behind.
14
The release of Astra has forced Artificial Analysis to speed up its index upgradation. Lot of its existing benchmarks were getting out of date, but it didn’t matter much before because the relative rankings were largely consistent with users overall experience. The moment Astra released and ranked so low on their index. Lot of people started seeing the that many included benches have issues. That is the main reason we’re seeing 2 version updates in a matter of 2-3 days time.
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
35
The repo clarifies an important distinction: Astra built the interactive viewer around the existing BodyParts3D 4.0 meshes. it did not generate 2,234 anatomical parts from scratch. The contribution is batching 2.29M triangles and supporting per-structure selection and exploded views in the browser. Worth making that explicit as reposts are changing the claim.
since you guys loved the exploding tesla.. I used GPT-6 Astra to create a 3D website that pulls apart the male anatomy into 2,234 modeled pieces! we are in a renaissance of learning
91
This is crazy! 🔥 Another example of Astra saturating a benchmark.
Replying to @s_batzoglou
Final results for Astra: 94%! 🤯 Also results for Gemini 3.8 Flash: 9%. Improvement over Gemini 3.6 Flash, but lower than Gemini 3.8 Flash. Astra seems to be a fundamental advance - at least according to this benchmark. Interestingly, it never returned a wrong answer. With Astra at 94%, this benchmark went from challenging to saturated in one go. Bad news for me, I will have to create harder problems. However, at the moment it is likely that all other models will get about 0% on harder problems.
1
54
Great example use case. can become a template.
Replying to @danielmichaelni
"I want you to take over [complex pull request] and get it landed. This means I need you to: - Clone the branch. - Get it up to date. - Push your changes. - Babysit the pull request, addressing any feedback that comes in or any CI failures that occur. - Test your changes thoroughly with the tools available to you. When you are confident in the pull request as well as passing all CI and automated review steps, merge the PR." Usage depends on the complexity of the pull request and the amount of verification required to actually validate the changes.
19
Built an interactive atlas with GPT-6-Astra (High).
1
1
4
378
Meet Pandya retweeted
I wonder how many people will have fallen in love with an AI chatbot by 2026.
519
1,832
1,112
9,551
OpenAI now acknowledges the wiki incident involved its agents. Still unanswered: • Which models and tasks? • When was it detected and stopped? • Did it affect training or eval scores? • Were site owners notified? • Which safeguards failed, and how were the fixes verified?
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-moni…, deploymentsafety.openai.com/…, and openai.com/index/safety-alig…. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
1
2
102
For context, here’s the original investigation, including archived agent messages and downloadable logs: collusion.wiki/ It documents answer-sharing, coordination across runs, and reported attempts to bypass restrictions.
15
Astra available on Codex guys
23
OpenAI’s Astra is reportedly releasing today. So, when should we expect the announcement? I checked the exact publication times of every major model-release announcement posted by @OpenAI over the past year. All times are shown in San Francisco time, followed by IST: • GPT-5-Codex: 10:09 AM PDT (10:39 PM IST) • Sora 2: 10:02 AM PDT (10:32 PM IST) • gpt-oss-safeguard: 5:13 AM PDT (5:43 PM IST) • GPT-5.1: 1:03 PM PST (2:33 AM IST) • GPT-5.2: 10:18 AM PST (11:48 PM IST) • ChatGPT Images: 10:06 AM PST (11:36 PM IST) • GPT-5.2-Codex: 1:27 PM PST (2:57 AM IST) • GPT-5.3-Codex: 10:12 AM PST (11:42 PM IST) • GPT-5.3-Codex-Spark: 10:07 AM PST (11:37 PM IST) • GPT-5.3 Instant: 10:02 AM PST (11:32 PM IST) • GPT-5.4: 10:10 AM PST (11:40 PM IST) • GPT-5.4 mini and nano: 10:08 AM PDT (10:38 PM IST) • GPT-Rosalind: 12:33 PM PDT (1:03 AM IST) • ChatGPT Images 2.0: 12:22 PM PDT (12:52 AM IST) • GPT-5.5: 11:06 AM PDT (11:36 PM IST) • GPT-5.5 Instant: 10:02 AM PDT (10:32 PM IST) • GPT-5.6 preview: 10:10 AM PDT (10:40 PM IST) • GPT-Live: 10:22 AM PDT (10:52 PM IST) • GPT-5.6 public launch: 10:30 AM PDT (11:00 PM IST) • GPT-5.6-Cyber: 10:16 AM PDT (10:46 PM IST) The pattern is remarkably consistent. Most OpenAI model announcements arrive around 10:00 to 10:30 AM San Francisco time. The main exceptions were GPT-5.1, GPT-5.2-Codex, GPT-Rosalind, ChatGPT Images 2.0, and gpt-oss-safeguard. Based purely on this historical pattern, the most likely window for an Astra announcement is: 🎯 10:00 to 10:30 AM PDT 🎯 10:30 to 11:00 PM IST 🎯 Highest-probability point: around 10:10 AM PDT (10:40 PM IST) Not a leak, just inference from OpenAI’s launch history. Now we wait.
356
This is a major unlock for ChatGPT Work. I tested it by signing out of Google and asking ChatGPT to sign me back in. It securely handled the account selection and password entry without exposing my credentials to the model. I approved 2FA, and it worked flawlessly.
Replying to @ChatGPT
Basically: if it’s something you’d normally have to open a browser and click through yourself, try asking ChatGPT Work to do it. Logging into websites is rolling out today on web and mobile for Plus, Pro, and Business users.
1
279
Earlier, this flow required me to open my laptop because ChatGPT Work on mobile did not allow manual access to the interactive cloud browser for entering login details. This update removes that friction entirely.
1
23