🇰🇷 Engineer at @realAsteromorph | F1 enthusiast | Prev. Tech Lead @hashed_official | @kaggle Notebook Grandmaster
Seoul, Korea
Joined December 2015
- Tweets3.7K
- Following3.7K
- Followers26K
- Likes25K
Personal update: after four years as Data/Tech Lead at @hashed_official, I'm joining @realAsteromorph as a Research Engineer, working on AI for Science.
Hashed was the best team I've been part of. Working at the intersection of business and tech, those four years genuinely widened my perspective and my capabilities. @simonkim_nft , @Hashed_HongPro , @0xryankim @baekkyoumkim, thank you for trusting me with so much. 🙏
And if a founder ever asks me what Hashed is like from the inside, my answer is simple: I've never seen a team care this much about its people and its founders.
Asteromorph is one of the youngest and most talented teams I've come across, taking on how AI can accelerate scientific discovery. I'm looking forward to the days ahead.
I won't be at the ICML main conference, but with the whole field gathering in Seoul this month, I'll be around main workshops, the side events and meetups. If you're in town, I'd love to meet and talk AI for Science.🚀
benchmark saturation (or benchmaxxing) is happening faster than before, but new benchmarks are appearing just as fast and i don't think they're being validated enough.
the people who can make hard, important problems and good data are only going to matter more.
a lot of people now try the models themselves instead of just comparing benchmark scores. that's what crossing the chasm feels like.
Agent benchmark evaluation will become an increasingly important task.
And when thinking about "what makes a good benchmark," the 3-axes discussed in @SnorkelAI's recent presentation are a really useful framework:
- Environment Complexity
- Autonomous Horizon
- Output Complexity
I’ve also been wondering whether there should be a "time-related axis", covering things like contamination and how well a problem remains relevant over time, and when we might be able to detect those signals.
@fredsala's talk and a few words of advice were helpful in thinking through this, and gave me more to think about :)
We are already using AI on problems that are too complex, or too large, for even a small group of people to fully understand and verify.
This changes what evaluation means. As benchmarks include more unseen tasks, false positives and false negatives become harder to avoid. A model may pass for the wrong reason, or fail for a reason the rubric does not capture.
So the important question is not just whether the model got the answer right. It is whether the rubric is the right way to judge the answer at all. And when the model is wrong, we need to ask what kind of wrong it is. Why did it fail? How did it fail? What does the failure reveal?
This is why expertise does not go to zero.
As AI gets better, the need for people who can interpret, validate, and reason about its outputs becomes more important, not less.
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
openai.com/index/separating-…
ICML week, but still building at night
a readable, searchable index of scientific llm benchmarks.
> countless benchmarks and good scores, but still so few meaningful discoveries. why?
What does it mean for an AI model to be “good at science”?
I started collecting Scientific LLM benchmarks in one place - across math, physics, chemistry, materials science, biology, and scientific agents.
github.com/subinium/Awesome-…
Suggestions are welcome.
What does it mean for an AI model to be “good at science”?
I started collecting Scientific LLM benchmarks in one place - across math, physics, chemistry, materials science, biology, and scientific agents.
github.com/subinium/Awesome-…
Suggestions are welcome.
github is genuinely a great dev platform for agents.
if you already get why plan mode matters, then getting fluent with gh cli (issues, milestones, tags, releases) will easily push your productivity 10x beyond where it was.
with 4.7 or 5.5 or both, give a senior dev unlimited tokens and unlimited compute, and they can now ship 100+ PRs across 10+ open source repos every single day. And this is only going to accelerate.
everything is collapsing into a question of money.