For people who do lots of LLM as judge stuff, is Sol too nitpicky or is Fable too lax? They disagree so much in my recent experiments
Why the fuck is the only good Caltrain/BART interchange at Millbrae? You cannot pretend to be a real city and do this
It's so over
Replying to @wtgowers
AI has now solved a major open problem -- one of the best known Erdos problems called the unit distance problem, one of Erdos's favourite questions and one that many mathematicians had tried.
openai.com/index/model-dispr…
Charlie London retweeted
Given black-box access to a Transformer's output, can we efficiently recover its parameters?
We analyse the learnability of attention-based models with query access in our new work. Accepted at #ICML2026 🎉
Work done with @shahkulin98, @mhahn29 and Varun Kanade.
🧵
If enough people sign up I'll do it as bane
If models can think for 100,000 tokens, why do they still lose the plot?
Come join us for this AI4Science on alphaXiv talk: Long-Horizon Reasoning in LLMs.
In this session, Sumeet Motwani (@sumeetrm) and Charles London (@CharlieLondon02) will share recent work on both training and evaluating models that can reason over much longer chains of thought.
Their LongCoT benchmark tests whether models can handle long chains of dependent reasoning across different fields. Each step is solvable on its own, but the full problem requires planning, state tracking, backtracking, and avoiding compounding errors. Even the best models still score below 10%.
They will also discuss h1, which trains long-horizon reasoning by chaining short problems into longer dependency graphs, then using RL with outcome-only rewards and a gradually harder curriculum.
So if longer context windows are not enough, what does it actually take to make models reason reliably over long scientific and technical workflows?
Whether you’re working on frontier LLMs, AI4Science, reasoning, or just curious about what current models still cannot do, you should definitely check this talk out!
🗓 Friday May 15th 2026 · 11 AM PT
🎙 Featuring Sumeet Motwani and Charles London
💬 Casual Talk + Open Discussion
Charlie London retweeted
Two British people self-isolating at home in UK after potential exposure to hantavirus on cruise ship, UKHSA says bbc.in/4eCLI5z
h1 and LongCoT are going to be at ICML! See you in Seoul 🇰🇷
Love for all my coauthors, but particularly my goat Sumeet
🚨How do we improve long-horizon reasoning capabilities by scaling RL with only existing data?
Introducing our new paper: "h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning"🫡
> RL on existing datasets saturates very quickly
> Reasoning over complex interdependent problems is incredibly important, but we currently lack enough long-horizon reasoning data
> Long-horizon problems are hard, which means training signal is sparse. We’d need a way to provide dense supervision
Our solution composes existing short-horizon data to form a synthetic curriculum that keeps growing in complexity! This allows us to scale RL on the same dataset while avoiding saturation, with curriculum acting as dense rewards.
At a small scale, we see massive in-domain long-horizon improvements, which transfer to significantly harder benchmarks. Training on composed 6th grade math problems leads to strong gains on AIME! 1/N🤿🧵
Charlie London retweeted
Muon leads to severely miscalibrated models!
This is just one of the results of this new paper of ours:
In “Too Sharp, Too Sure” we show calibration error tracks loss curvature during training and we tie both to margin tails.
There's great info about how we might train rlm's to solve difficult problems here, particularly around prompting and data filtering to generate good traces. Thanks to @chenxiao_yang_, we already know they're theoretically very powerful, it's now about unlocking it
New mini experiment + blogpost + trajectories!
tldr; we boost performance of RLM(GPT-5.2) to double the best performing number (38.7% --> 65.6%) on LongCoT-mini without any training! An example of the mismanaged geniuses hypothesis (MGH) we (@zli11010, @lateinteraction) proposed earlier this month.
The LongCoT benchmark showed that frontier LMs and RLMs struggled to solve difficult compositional reasoning tasks. The paper generally attributes this to the RLMs inability to perform task decomposition, but we argue this is more our fault in how we prompt them; this capability is fully available to GPT-5.2 with an RLM harness!
Building on @raw_works's insightful blogpost and @sumeetrm / @CharlieLondon02 et al.'s incredibly useful benchmark, where they originally found RLMs to be incapable of solving the MATH and CS splits altogether. We did not train anything since the release of the initial benchmark.
To be fully transparent, these results are not meant to be added to their leaderboard either; benchmarks measure isolated capabilities, and we focus on showing (through different, rather specific prompting) that the capabilities required to solve these tasks are available to the models without additional training! It also has implications about how we would go about training these systems. Full blog below, it's a nice read :)
Great work with rlm's on longcot from Alex!
New mini experiment + blogpost + trajectories!
tldr; we boost performance of RLM(GPT-5.2) to double the best performing number (38.7% --> 65.6%) on LongCoT-mini without any training! An example of the mismanaged geniuses hypothesis (MGH) we (@zli11010, @lateinteraction) proposed earlier this month.
The LongCoT benchmark showed that frontier LMs and RLMs struggled to solve difficult compositional reasoning tasks. The paper generally attributes this to the RLMs inability to perform task decomposition, but we argue this is more our fault in how we prompt them; this capability is fully available to GPT-5.2 with an RLM harness!
Building on @raw_works's insightful blogpost and @sumeetrm / @CharlieLondon02 et al.'s incredibly useful benchmark, where they originally found RLMs to be incapable of solving the MATH and CS splits altogether. We did not train anything since the release of the initial benchmark.
To be fully transparent, these results are not meant to be added to their leaderboard either; benchmarks measure isolated capabilities, and we focus on showing (through different, rather specific prompting) that the capabilities required to solve these tasks are available to the models without additional training! It also has implications about how we would go about training these systems. Full blog below, it's a nice read :)
Charlie London retweeted
LongCoT is adding two new leaderboards! Due to the interest in agents (particularly RLMs), we’re adding a “Restricted Harness” and an “Open Harness” leaderboard.
GPT 5.2 RLM from our paper is SOTA on “Open Harness” at 25.12%. We expect tool-use SOTA to exceed this very soon!
On “Open Harness”, we allow all tool-use and code execution. On “Restricted Harness”, models may manage context, call subagents, etc, but may not write specific solver code (e.g. writing a BlocksWorld or Sudoku solver). We’re particularly excited about this leaderboard, as it allows agents to do their own context management, while sticking to LongCoT’s goal of testing models’ intrinsic reasoning capabilities.
Charlie London retweeted
We’re releasing LongCoT, an incredibly hard benchmark to measure long-horizon reasoning capabilities over tens to hundreds of thousands of tokens.
LongCoT consists of 2.5K questions across chemistry, math, chess, logic, and computer science. Frontier models score less than 10%🧵
Charlie London retweeted
The attention on LongCoT is great! It's far from solved (GPT 5.2 w/out tools gets 9.8%).
Out-of-the-box, a GPT 5.2 RLM gets 25% (see Figure 7). Better prompting/training should push RLMs past this.
Comparing RLMs to no-tool baselines? See our 🧵of tips
nitter.cf/sumeetrm/status/204480…
We've identified 21 unsolvable problems in LongCoT-mini maths (21 of 507 mini questions), traced to missing information in some subproblems. This error does not affect other questions. Affected question IDs: github.com/LongHorizonReason….
We're working on a fix, ships very soon!
I do wonder whether Anthropic's use of a new tokenizer is to improve downstream scaling performance. After reading the Cagnetta's "Deriving Neural Scaling Laws", I did wonder whether it would be possible to build a tokenizer that optimizes their gamma and beta params.