@Benjamin_eecs

Incoming PhD @StanfordNLP | BS @PKU1898 | Building continually self-improving AI | Prev @DeepSeek_AI @AIatMeta | DeepSeek-V1/V2/VL/Prover ALE SPIRAL SPICE SPADE

South Korea
Joined February 2022
Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
20
127
34
703
178,849
Bo Liu (Benjamin Liu) retweeted
Scientific discovery is one of the most exciting applications of LLMs, but we lack data & evals on the scientific process. ScholarCatalyst captures which prior papers can advance new research on a problem, rather than citations or topical relevance. Strikingly, even frontier models struggle to identify which prior research is helpful -- lots of room for future work! Paper, Code, Data: ohmyksh.github.io/project/Sc…
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
9
17
3
193
13,730
Bo Liu (Benjamin Liu) retweeted
We release a benchmark that checks if retrievers (embedding models & ai agents) can find prior works that "inspired" a research paper, extending beyond simply finding relevant papers. In my point of view, having a good model on this benchmark can enable assessing novelty of papers or generating novel research ideas. Checkout @ohmyksh's post!
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
4
9
34
2,094
Bo Liu (Benjamin Liu) retweeted
In here is a video of The Bitter Lesson as a pretty-good country-music song. Enjoy. oneusefulthing.org/p/the-dot…
7
24
5
232
15,860
Bo Liu (Benjamin Liu) retweeted
To tackle new scientific problems, we often get inspirations from prior work. Can Deep Research agents and search help us to identify such “catalyst” papers buried in literature? Our new benchmark, built with hundreds of scientists, shows substantial room for improvements.
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
4
10
1
83
6,743
Bo Liu (Benjamin Liu) retweeted
Chief in AI 100 (2026): Yejin Choi, Professor, Stanford (MacArthur Fellow); Distinguished Scientist, NVIDIA. Profile: aibuildersnetwork.org/chief-… Meet leaders like these at the conference, Oct 14–16: aibuildersnetwork.org/confer… #ChiefInAI100 #aibgc2026
2
2
1,994
Research taste will matter more and more as AI starts doing research on its own. Part of taste is knowing which old papers to build on. ScholarCatalyst is a first step toward measuring that automatically :))
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
2
7
42
2,405
Bo Liu (Benjamin Liu) retweeted
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
17
53
8
336
23,827
Bo Liu (Benjamin Liu) retweeted
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
6
56
13
258
39,567
Bo Liu (Benjamin Liu) retweeted
FrontierPhysics update #3 We're wrapping up the v0.1 leaderboard. From initial runs on 56 real physics research tasks, the strongest frontier model passes only 22% of its runs. How we got a number we trust 🧵
2
4
1
14
2,205
Bo Liu (Benjamin Liu) retweeted
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
42
185
29
1,408
275,236
Bo Liu (Benjamin Liu) retweeted
⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the role & impact of humans in the loop Blog post: facebookresearch.github.io/R… Key takeaways: 1) We find human-agent collaboration gives big wins over agents alone - fine-grained feedback in ideation stage crucial - autobench can be used to measure this in the future with stronger agents 2) We show that it's possible to make *autoresearch benchmarks* for AI research using this recipe – full recursive improvement loop! 3) Important Ingredients: Autobenchmark creation works best with feedback from two sources: benchmark solvers + external verifiers (human+AI). 🧵1/5
1
64
19
525
59,014
Bo Liu (Benjamin Liu) retweeted
Claim: we've solved the AI slop problem (!) 💩🧹✨ Blog post: facebookresearch.github.io/R… 🧵1/5 Key idea: take *expert* human writing and learn rubrics that find the gap between experts and models. Train with those rubrics. We train with RL-XAR (RL with eXpert Aligned Rubrics) & see large performance gains on writing scientific paper sections, Pulitzer prize novel continuations and high quality Wikipedia pages.
1
177
70
1,908
221,236
Bo Liu (Benjamin Liu) retweeted
🔥 Jev is on fire! 😎 So we ask Reef: /reefine add jev as a tool Reef then evolves the agent, and adds Jev into its harness. We use this evolved agent to hunt for papers related to "Self-evolving Agents" published in 2026 on arXiv. The agent found and screened 120 papers in 9 seconds! Come and use Reef to evolve your agent: github.com/Human-Agent-Socie…
7
26
5
61
104,405
Bo Liu (Benjamin Liu) retweeted
Reef just dropped v0.1.1 🚀 In the past week, Reef has: → 🌟 Crossed 5K GitHub stars, adding ~1.6K → 🔧 Contributions from 35 developers so far Try out v0.1.1 and tell us what you want your agent to keep learing. Checkout Github Repo at: github.com/Human-Agent-Socie…
1
34
4
135
38,168
Bo Liu (Benjamin Liu) retweeted
We built an all-synthetic simulation of a hospital that can be shared publicly without any issues. Our data quality is so high, physicians cannot tell the difference between real/synthetic patients. The data is fully verified -- perfect for RL.
24
85
9
1,141
142,104
Bo Liu (Benjamin Liu) retweeted
This is cool! We’ll have grown up LLMs nurturing baby LLMs to learn by self-play in the crib in no time! More seriously, it’s a good demonstration of how general inductive biases can be an effective “Universal Grammar” for more general learning about the world.
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
4
13
116
13,465
Bo Liu (Benjamin Liu) retweeted
Your agent can now grow new abilities, just by asking. Reef now supports personalized harness evolution with /reefine: describe what you want your agent to do. Reef builds the change, checks it, and ships a new version. Examples: 💬 "> /reefine add a /chat mode for faster responses" 🔊 "> /reefine tell me out loud when you're done" ⚡ "> /reefine add jev as a tool": 120 papers screened in 9 s Works with any agent harness through a Reef adapter: Codex, Hermes, OpenCode, Pi and Terminus 2 today. Try it here: github.com/Human-Agent-Socie… #agent #harness #rsi #llm
9
39
13
120
164,697
Bo Liu (Benjamin Liu) retweeted
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
53
146
31
1,303
298,591
DeepSeek will not announce «RSI», they'll just explain it in passing as section 6 of their paper on sandbox infrastructure
5
19
415
13,925
Bo Liu (Benjamin Liu) retweeted
Releasing our runtime dynamic compression framework that achieves 1.5-2.0 bit compression at high quality. This is integrated into the bitsandbytes2 library, which starts as a private beta today. Paper: timdettmers.com/papers/runti… Private beta signup: forms.gle/Pwtp8CVRULdoEnY79
20
48
3
294
31,274