@phianderssoni
iAccount based inSweden
About this account
- Account based in
- Sweden
- Connected via
- Sweden App Store
Account-level information from X, not a live location or the device used for a specific post.
Philip retweeted
openai.com/index/astra-for-l…
"Astra for Law combines GPT‑6 Astra with a powerful legal search index and instructions for legal analysis and writing.
Together, they amplify Astra’s capabilities across the legal practice, while giving firms and legal technology companies the freedom to build their own applications and workflows."
Philip retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
Philip retweeted
Big overhaul on DeepSeek V4.1 using an encoder-decoder setup.
Tbh they should have called it DeepSeek V5!
Super cool and refreshing, though!
Now available: ChatGPT for Financial Services.
This is a tailored ChatGPT Work experience that combines built-in financial data with GPT-6 Astra’s reasoning.
Teams can develop research, build financial models, and create customized client materials.
openai.com/index/introducing…
Philip retweeted
Hacker house leaders in Europe or the women powering the future of Nordics and Baltics
@Stockholmschool @KiuasHQ Ruum Tallinn
Philip retweeted
I've spent the last 1.5 years working with the best AI startups globally. This distills my learnings on prompts. Engineers & builders, I hope it's useful 🤝
Astra is fully rolled out to Plus, Pro, Business, and Enterprise users in Codex and ChatGPT Work.
Go build!
And if you need inspiration, watch Astra in action, live: openai.com/gpt-tv/
Philip retweeted
The OpenAI/Navier-Stokes mathematician beef is crazy.
Quick recap + my thoughts:
Two mathematicians, Tristan Buckmaster and Levent Alpöge (works at Anthropic), were pursuing the Navier-Stokes Millennium Prize Problem.
Their idea was to construct a perfectly smooth force that makes the fluid equations blow up in finite time.
They had already proven related results and stored all their unpublished drafts in private Codex sessions.
Then, boom, OpenAI drops this banger post, claiming 10,000 agents solved Navier-Stokes using what Buckmaster describes as a suspiciously similar approach.
Buckmaster hears about this and thinks:
“Did these mfs train on our Codex chats?”
He asked whether OpenAI’s model had accessed or been trained on their private sessions.
OpenAI says no specific user data was accessed. But it also says it cannot rule out de-identified data from their usage having helped improve the model.
So: clear denial of direct access, no clear answer on indirect training.
But it would make sense for OpenAI to train on Codex sessions: It's probably one of the main reasons they are letting people go crazy with the resets:
Infinite high quality agentic training data hack
This whole situation reminds me of the Claude Mythos finding a real FreeBSD vulnerability earlier this year.
It was later revealed that the exact same bug had been found and patched in closely related MIT Kerberos code 19 years earlier.
...So it was in the training data.
FreeBSD vulnerability solution also leveraged thousands of agents over several hours... So maybe that just increases the changes of one agent getting lucky with it's weights?
Get what I'm saying?
The same question now hangs over OpenAI.
If the approach came from the mathematicians’ work, this is still an insane demonstration of what AI can learn and execute, but also a massive question of consent and credit.
If Buckmaster and Alpöge's work isn't in the data, and 10,000 agents independently discovered the same route from scratch…
This might just be AGI
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Philip retweeted
Replying to @ChaseLochmiller @OpenAI
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
ChatGPT Images 2.5—faster, sharper, smarter, with better tools for creating whatever you can dream of.
- Faster image generation to keep your ideas flowing
- Improved fidelity for more natural, recognizable images
- Consistent details across multiple edits
- Comment-based edits to change only what you want
Philip retweeted
We want to celebrate with people using GPT-6. We did this for GPT-5.5 and it was really fun.
We’re getting together in SF on September 16 to talk about the model, what we should build next, and mostly just to hang out.
Apply by Sep 10:
gpt6-launch-event.openai.cha…
Philip retweeted
GPT-6 Astra is almost at human performance in 3D spatial understanding.
It ranks #1 on Blueprint-Bench 2, a benchmark where AI agents draw floorplans from photographs of apartment interiors.
Philip retweeted
OpenAI DevDay is going global.
Starting this October, DevDay Exchange is bringing builders together to swap build notes, share real projects, and meet the teams building OpenAI tools in:
Bengaluru
Tokyo
Seoul
Berlin
Paris
London
São Paulo
Mexico City
We're ready to see what developers around the 🌎 are pushing to prod.
Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)!
It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing.
GPT-6 Astra by @OpenAI takes the top spot in Code Arena: WebDev with a score of 1797 pts. This opens up a solid +35pt lead over #2 Claude Fable 5.1 (Max) at 1762 pts and #3 Claude Opus 5 (Max) at 1688 pts.
This is a significant improvement from GPT-5.6 Sol (xHigh) at +180 pts, ranked at #13.
Category level votes still incoming, but already we see it at #1 in: Data & Analytics, Consumer Product, Content Creation Tools and #2 in Gaming and Simulations.
Stay tuned for other categories like Brand & Marketing, Reference-Based Design and Full Stack rankings.
Congrats to the @OpenAI team on this release!
Philip retweeted
HUGE! You can now use voice mode with *any* codex thread - even existing ones!
Quality of Life improvement!!
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0.
GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
with GPT 6 astra you can feel the agi.
- Codex is the only interface i use.
- browser use and CUA are fast and precise. the browser is now just an embedded tool.
- the model has actual visual taste. the output finally looks designed instead of generated.
- CAD and 3D crossed a threshold. astra can build a house in blender and turn it into a walkable unreal engine world.
- long running work finally holds together. less babysitting, fewer resets, more finished work.
- during cyber testing it discovered and exploited two previously unknown zero days.
my take: software is something agents use, not humans.
we live in the future.