@MatternJustusi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Germany App Store
Account-level information from X, not a live location or the device used for a specific post.
Co-Founder @ProximalHQ | prev. research @PrimeIntellect, @MPI_IS and built revideo
San Francisco, CA
Joined March 2021
- Tweets1.2K
- Following909
- Followers8.5K
- Likes3.4K
Pinned Tweet
FrontierSWE v2 shows strong gaps between all frontier models that are not visible in other benchmarks
Fable 5.1 and also Opus 5 are far ahead of other models - it is also notable that GLM-5.3, the strongest open source model, is at #3 despite not having vision capabilities
Every Anthropic model released since Fable 5 has become more cost-efficient on FrontierSWE v2 while achieving higher scores.
Opus 5.5 is particularly cheap and fast. This matches the team's vibes when using it!
Replying to @ProximalHQ
Opus 5.5 is both the cheapest and highest-scoring model released by Anthropic
This trend has been constant since Fable 5 - every model released after it scored higher while being cheaper than its predecessor
Great to see FrontierSWE v2 on the Specialized Intelligence Index!
We partnered with @FireworksAI_HQ to bring FrontierSWE v2 to the Specialized Intelligence Index
Evals are essential for improving and safely deploying AI. We're excited to support the initiative!
We are hiring a generalist intern this fall to work closely with me on applied research!
You should be technical, but you will not be asked to write code all day - instead, we will work together across product, operations and ensuring our customers are happy
DMs open!
Astra is a very strong model!
I was the most surprised by its solution in Kolmogorov Audio Compression: Instead of writing a compression algorithm, Astra figured out that we had synthesized the audio programmatically and simply reverse engineered the code for it. This way, it was able to compress 600MB+ of audio data into 20KB
Very excited for this! Congrats to the team, cannot wait for new open models
Today, we are announcing our Series B funding round, valuing the company at more than $1B.
This round accelerates our next-gen Trinity models across diverse infrastructure, expands our work with the DOE and national labs on Genesis-Science-1, and enables us to build the platform teams need to build, evaluate, deploy, and operate open models in production.
We are grateful to our team, partners, open-source community, and investors.
Led by @Vista_Equity, Cambium Capital, and @emergencecap, with participation from AI10 Ventures, @Hitachi, IAG, @M12vc, @p7ventures, and @Wipro.
Building an autonomous lab from scratch to collect post-training data is an incredibly ambitious bet, but also the logical conclusion if you believe in RL
Extremely excited about the work Periodic is doing!
Interesting to see @deepseek_ai use FrontierSWE to study their multi-agent harness
I suspect that the score gap is smaller compared to Programbench since you benefit most from multi-agent setups in pure implementation tasks where difficulty comes from thoroughness and the ability to write a lot of code.
Performance engineering and other tasks that are more autoresearch-like seem to benefit less from parallelization
Frontier labs getting serious about safety also means evaluating which data vendors to work with, and being willing to pay a premium for those that actually care about quality.
I'm glad to see that most labs are doing so!
New research: Training a Misaligned Reward Seeker
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.
In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
Read more: alignment.anthropic.com/2026…
Justus Mattern retweeted
We need better evals.
One of the motivations behind working on FrontierSWE and releasing it was that at the time we were seeing many blogs about how agents running for many hours could accomplish incredible tasks like how:
- Cursor built a web browser from scratch
- Bun had an agent rewrite Bun in rust
- Anthropic built a C compiler.
We saw that many benchmarks in coding were saturated (SWE-Bench Pro etc) and wanted a benchmark that could accurately reflect ultra long horizon and extremely difficult software engineering tasks that would usually take humans weeks to months in an attempt to measure AI progress
With Muse Spark 1.3 and Gemini Flash 3.8 scoring similarly as a Fable 5.1 or Astra on benchmarks like DeepSWE, its clear that the need for really good benchmarks that accurately reflect model progress that are both fair and hard for frontier models is much needed
I might be biased but FrontierSWE is the only coding benchmark that still remains unsaturated and most closely reflects the preferences of day to day developers
At Proximal, we are committing millions of dollars in the next few months to creating better evals for the community in domains like software engineering, science, hardware engineering, etc.
If you are excited about this kind of work, we'd love to chat with you. The team is very lean (3 people) and has a lot of autonomy and resources.
I don't think that GLM-5.3 or Kimi K3 are benchmaxxed. The models are genuinely competitive with or better than models that are not Astra or Fable
Eval forensics: No strong signal of GLM or Kimi benchmaxxing?
Most of v2 did not exist when GLM-5.3 and Kimi K3 were trained. v1’s 17 tasks went public in April. K3 shipped July 16, GLM-5.3 August 14. v2 landed September 2–3 with 21 new tasks. Those 21 were not in the training window. If GLM or K3 had maxed the public set, they would spike on the old tasks and flatten on the new ones.
More details from Grok: nitter.cf/i/grok/share/dd5aa9a87…
Justus Mattern retweeted
Read the thread. This is a really REALLY nice eval
Justus Mattern retweeted
reading traces is so fun
(the prompt does not talk about littlefs, the model inferred it from the tests + pre-train knowledge)
The results that led to our evaluation harness is imo some of the most interesting work that went into FrontierSWE v2
Simply reminding GPT-5.6 Sol that it has a lot of time left boosts its scores by a lot. Hoping to see other long-horizon benchmarks adopt this methodology!
Replying to @ProximalHQ
FrontierSWE v2 uses proximus, our minimal agent harness which extends mini-swe-agent for ultra-long horizon tasks.
When an agent wants to submit a solution, proximus encourages it to work longer. This way, models don't submit prematurely and achieve much better results
I'm extremely proud of the care that went into this release. The effort it takes to build good benchmarks is huge, and the team did an incredible job here
We'll follow up with results for Gemini 3.8, GLM-5.3 Flash and other models soon. Please let us know if you have feedback!
People with both research taste as well as good commercial instincts are true unicorns.
We are hiring someone to help build our applied research function and work closely with customers. If you are an engineer or researcher that wants to learn these skills, please reach out!
Justus Mattern retweeted
Fable 5.1 is the strongest model on FrontierSWE v2 (still unreleased) with a score of 0.57, beating Opus 5 (0.52), Fable 5 (0.48) and GPT-5.6 Sol (0.32)
FrontierSWE v2 will be released this week!
Fable 5.1 is the strongest model on the (still unreleased) FrontierSWE v2 with a score of 0.57, beating Opus 5 (0.52), Fable 5 (0.48) and GPT-5.6 Sol (0.32)
We will release the full benchmark this week!
The disastrous consequences of training on hackable environments is one of the reasons why I don't believe in a fragmented market of small, low-quality data vendors
QA that is actually good is difficult to build and will only become harder as models get better
We’re sharing an update on our alignment and security efforts.
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe:
1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards
2. An update on our alignment assessment
3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them
4. How we hardened our security practices earlier this year to prepare for Mythos-class models
Read more: anthropic.com/news/improving…