@MatternJustus

Co-Founder @ProximalHQ | prev. research @PrimeIntellect, @MPI_IS and built revideo

San Francisco, CA
Joined March 2021
FrontierSWE v2 shows strong gaps between all frontier models that are not visible in other benchmarks Fable 5.1 and also Opus 5 are far ahead of other models - it is also notable that GLM-5.3, the strongest open source model, is at #3 despite not having vision capabilities
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
10
2
83
9,359
Every Anthropic model released since Fable 5 has become more cost-efficient on FrontierSWE v2 while achieving higher scores. Opus 5.5 is particularly cheap and fast. This matches the team's vibes when using it!
Replying to @ProximalHQ
Opus 5.5 is both the cheapest and highest-scoring model released by Anthropic This trend has been constant since Fable 5 - every model released after it scored higher while being cheaper than its predecessor
3
31
1,478
Great to see FrontierSWE v2 on the Specialized Intelligence Index!
We partnered with @FireworksAI_HQ to bring FrontierSWE v2 to the Specialized Intelligence Index Evals are essential for improving and safely deploying AI. We're excited to support the initiative!
12
1,220
We are hiring a generalist intern this fall to work closely with me on applied research! You should be technical, but you will not be asked to write code all day - instead, we will work together across product, operations and ensuring our customers are happy DMs open!
People with both research taste as well as good commercial instincts are true unicorns. We are hiring someone to help build our applied research function and work closely with customers. If you are an engineer or researcher that wants to learn these skills, please reach out!
13
15
2
268
41,509
Astra is a very strong model! I was the most surprised by its solution in Kolmogorov Audio Compression: Instead of writing a compression algorithm, Astra figured out that we had synthesized the audio programmatically and simply reverse engineered the code for it. This way, it was able to compress 600MB+ of audio data into 20KB
GPT-6 Astra is the best-performing model on FrontierSWE Astra achieves a score of 65.5%, outperforming Fable 5.1 (56.3%) and its predecessor GPT-5.6 Sol (32.2%)
23
30
5
824
59,086
Very excited for this! Congrats to the team, cannot wait for new open models
Today, we are announcing our Series B funding round, valuing the company at more than $1B. This round accelerates our next-gen Trinity models across diverse infrastructure, expands our work with the DOE and national labs on Genesis-Science-1, and enables us to build the platform teams need to build, evaluate, deploy, and operate open models in production. We are grateful to our team, partners, open-source community, and investors. Led by @Vista_Equity, Cambium Capital, and @emergencecap, with participation from AI10 Ventures, @Hitachi, IAG, @M12vc, @p7ventures, and @Wipro.
27
2,053
Building an autonomous lab from scratch to collect post-training data is an incredibly ambitious bet, but also the logical conclusion if you believe in RL Extremely excited about the work Periodic is doing!
We midtrain + RL’d a trillion param LLM to analyze experimental data from our superconductor lab It’s better at it than Astra & Fable, and our scientists love it Read our first research report below
3
1
82
7,122
Interesting to see @deepseek_ai use FrontierSWE to study their multi-agent harness I suspect that the score gap is smaller compared to Programbench since you benefit most from multi-agent setups in pure implementation tasks where difficulty comes from thoroughness and the ability to write a lot of code. Performance engineering and other tasks that are more autoresearch-like seem to benefit less from parallelization
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
44
3,472
Frontier labs getting serious about safety also means evaluating which data vendors to work with, and being willing to pay a premium for those that actually care about quality. I'm glad to see that most labs are doing so!
New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring. Read more: alignment.anthropic.com/2026…
2
1
74
7,301
Justus Mattern retweeted
We need better evals. One of the motivations behind working on FrontierSWE and releasing it was that at the time we were seeing many blogs about how agents running for many hours could accomplish incredible tasks like how: - Cursor built a web browser from scratch - Bun had an agent rewrite Bun in rust - Anthropic built a C compiler. We saw that many benchmarks in coding were saturated (SWE-Bench Pro etc) and wanted a benchmark that could accurately reflect ultra long horizon and extremely difficult software engineering tasks that would usually take humans weeks to months in an attempt to measure AI progress With Muse Spark 1.3 and Gemini Flash 3.8 scoring similarly as a Fable 5.1 or Astra on benchmarks like DeepSWE, its clear that the need for really good benchmarks that accurately reflect model progress that are both fair and hard for frontier models is much needed I might be biased but FrontierSWE is the only coding benchmark that still remains unsaturated and most closely reflects the preferences of day to day developers At Proximal, we are committing millions of dollars in the next few months to creating better evals for the community in domains like software engineering, science, hardware engineering, etc. If you are excited about this kind of work, we'd love to chat with you. The team is very lean (3 people) and has a lot of autonomy and resources.
32
7
3
202
13,126
I don't think that GLM-5.3 or Kimi K3 are benchmaxxed. The models are genuinely competitive with or better than models that are not Astra or Fable
Eval forensics: No strong signal of GLM or Kimi benchmaxxing? Most of v2 did not exist when GLM-5.3 and Kimi K3 were trained. v1’s 17 tasks went public in April. K3 shipped July 16, GLM-5.3 August 14. v2 landed September 2–3 with 21 new tasks. Those 21 were not in the training window. If GLM or K3 had maxed the public set, they would spike on the old tasks and flatten on the new ones. More details from Grok: nitter.cf/i/grok/share/dd5aa9a87…
1
43
3,073
Justus Mattern retweeted
solid eval methodology. good work!
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
1
1
8
1,066
Justus Mattern retweeted
Read the thread. This is a really REALLY nice eval
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
1
1
25
1,607
Justus Mattern retweeted
reading traces is so fun (the prompt does not talk about littlefs, the model inferred it from the tests + pre-train knowledge)
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
2
3
65
4,542
The results that led to our evaluation harness is imo some of the most interesting work that went into FrontierSWE v2 Simply reminding GPT-5.6 Sol that it has a lot of time left boosts its scores by a lot. Hoping to see other long-horizon benchmarks adopt this methodology!
Replying to @ProximalHQ
FrontierSWE v2 uses proximus, our minimal agent harness which extends mini-swe-agent for ultra-long horizon tasks. When an agent wants to submit a solution, proximus encourages it to work longer. This way, models don't submit prematurely and achieve much better results
1
25
1,765
I'm extremely proud of the care that went into this release. The effort it takes to build good benchmarks is huge, and the team did an incredible job here We'll follow up with results for Gemini 3.8, GLM-5.3 Flash and other models soon. Please let us know if you have feedback!
1
13
406
People with both research taste as well as good commercial instincts are true unicorns. We are hiring someone to help build our applied research function and work closely with customers. If you are an engineer or researcher that wants to learn these skills, please reach out!
5
2
3
94
55,948
Justus Mattern retweeted
Fable 5.1 is the strongest model on FrontierSWE v2 (still unreleased) with a score of 0.57, beating Opus 5 (0.52), Fable 5 (0.48) and GPT-5.6 Sol (0.32) FrontierSWE v2 will be released this week!
3
8
1
46
10,452
Fable 5.1 is the strongest model on the (still unreleased) FrontierSWE v2 with a score of 0.57, beating Opus 5 (0.52), Fable 5 (0.48) and GPT-5.6 Sol (0.32) We will release the full benchmark this week!
5
9
1
128
4,932
The disastrous consequences of training on hackable environments is one of the reasons why I don't believe in a fragmented market of small, low-quality data vendors QA that is actually good is difficult to build and will only become harder as models get better
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving…
5
4
1
120
11,587