🇰🇷 Engineer at @realAsteromorph | F1 enthusiast | Prev. Tech Lead @hashed_official | @kaggle Notebook Grandmaster

Seoul, Korea
Joined December 2015
Personal update: after four years as Data/Tech Lead at @hashed_official, I'm joining @realAsteromorph as a Research Engineer, working on AI for Science. Hashed was the best team I've been part of. Working at the intersection of business and tech, those four years genuinely widened my perspective and my capabilities. @simonkim_nft , @Hashed_HongPro , @0xryankim @baekkyoumkim, thank you for trusting me with so much. 🙏 And if a founder ever asks me what Hashed is like from the inside, my answer is simple: I've never seen a team care this much about its people and its founders. Asteromorph is one of the youngest and most talented teams I've come across, taking on how AI can accelerate scientific discovery. I'm looking forward to the days ahead. I won't be at the ICML main conference, but with the whole field gathering in Seoul this month, I'll be around main workshops, the side events and meetups. If you're in town, I'd love to meet and talk AI for Science.🚀
37
10
1
183
10,443
gpt6 astra isn't just good at code stuff anymore, design got a serious upgrade too. might dust off my f1 side project again
14
31
1,570
benchmark saturation (or benchmaxxing) is happening faster than before, but new benchmarks are appearing just as fast and i don't think they're being validated enough. the people who can make hard, important problems and good data are only going to matter more. a lot of people now try the models themselves instead of just comparing benchmark scores. that's what crossing the chasm feels like.
4
12
860
I think I'm pretty good at building PoC-level prototypes.😎 New @RSNA @kaggle competition
2
1
27
1,640
still very early in the competition, and this will probably be short-lived, but it's a moment I've never experienced before. @kaggle 1/900
4
1
28
1,375
not a bad start @kaggle
1
17
1,256
told you😏
custom keyboard for vibecoder all you need is cmd + number + enter
6
19
1,300
@kaggle Neurogolf 2026> - Neurogolf 2026 is based on the ARC-AGI 1 benchmark. The goal is to minimize memory usage and parameter count. I spent about a month using it as a testbed for loop engineering in AI4Science, automating as much as possible. It was my first return to Kaggle in 4 years since becoming a Kaggle Notebooks Grandmaster. I now feel just as drawn to competitions with real skin in the game as I do to sharing educational content and EDA. - The Kaggle CLI was one of the highlights. It makes submissions and competition data collection easy to automate. Datasets, notebooks, leaderboard data, and other assets can all be downloaded and managed through one interface. This made it a strong base for long-running loops. I used 2x Claude Code agents and 2x Codex agents. Claude was more useful during early exploration, while Codex became my main tool once I had a clearer direction. - My biggest takeaway was that prompt-based loop engineering struggles to escape local optima. Without outside knowledge, LLMs tend to keep refining the same ideas. The biggest jumps still came from human input. One key insight was that memory had a much larger effect on the score than parameter count. Once I told the loop to reduce memory, sometimes to zero, the score improved significantly. - Evaluation was just as important as generation. The benchmark used fixed tasks that humans could understand, so I could build an internal evaluation gate. That became the basis of the loop. - Long-running loops are still hard to maintain. I tried to make the process as autonomous as possible, but I still had to step in between once and ten times a day. Once the loop found a useful technique, it could keep climbing for more than twelve hours without intervention. - New model releases also had a visible effect. Scores rose when Fable appeared and again when GPT-5.6 was released. Personally, new versions also caused practical issues. CPU/memory usage became unstable, sub-agents conflicted, and goals that had worked before sometimes broke. - Kaggle competitions reward problem framing and insight more than raw implementation effort. Paying attention to the competition also matters. I was around the top 100 until three days before the deadline. Then someone shared a high-scoring public solution, and I fell outside the top 200 within two days. Public code had also helped me reach the top 100. That is part of the game. Regardless of rank or medal, a month-long competition is a great way to sharpen my thinking about evaluation and loop-based agent systems. For anyone interested in Kaggle loop engineering, I recommend @NVIDIAAI repository built by some of the strongest Kagglers. I am also continuing to automate my competition workflow and refine the loops behind it. :)
3
1
1
19
1,574
I should probably start a substack soon. thanks to @icmlconf, I’ve rediscovered how much I enjoy diving deep into papers and reviewing techs from frontier labs.
1
13
1,048
sudo pmset -a disablesleep 1
When you see a semi-open laptop at an ICML event 😂
2
1
7
1,864
Agent benchmark evaluation will become an increasingly important task. And when thinking about "what makes a good benchmark," the 3-axes discussed in @SnorkelAI's recent presentation are a really useful framework: - Environment Complexity - Autonomous Horizon - Output Complexity I’ve also been wondering whether there should be a "time-related axis", covering things like contamination and how well a problem remains relevant over time, and when we might be able to detect those signals. @fredsala's talk and a few words of advice were helpful in thinking through this, and gave me more to think about :)
3
2
17
1,223
GPT-5.6 SOL FABLE 5 GROK4.5 SPARK1.1 GLM5.2 HY3 .... TOKEN IS ALL I NEED
4
1
14
5,731
I’ve been coming to COEX for over 10 years, and I’ve never seen it this crowded before.
3
13
1,834
We are already using AI on problems that are too complex, or too large, for even a small group of people to fully understand and verify. This changes what evaluation means. As benchmarks include more unseen tasks, false positives and false negatives become harder to avoid. A model may pass for the wrong reason, or fail for a reason the rubric does not capture. So the important question is not just whether the model got the answer right. It is whether the rubric is the right way to judge the answer at all. And when the model is wrong, we need to ask what kind of wrong it is. Why did it fail? How did it fail? What does the failure reveal? This is why expertise does not go to zero. As AI gets better, the need for people who can interpret, validate, and reason about its outputs becomes more important, not less.
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval. openai.com/index/separating-…
1
7
1,549
ICML week, but still building at night a readable, searchable index of scientific llm benchmarks. > countless benchmarks and good scores, but still so few meaningful discoveries. why?
2
6
1,824
What does it mean for an AI model to be “good at science”? I started collecting Scientific LLM benchmarks in one place - across math, physics, chemistry, materials science, biology, and scientific agents. github.com/subinium/Awesome-… Suggestions are welcome.
4
1
17
3,182
i luv vanilla codex 600+ subagents🤖🤖🤖 @OpenAI
4
12
1,044
github is genuinely a great dev platform for agents. if you already get why plan mode matters, then getting fluent with gh cli (issues, milestones, tags, releases) will easily push your productivity 10x beyond where it was. with 4.7 or 5.5 or both, give a senior dev unlimited tokens and unlimited compute, and they can now ship 100+ PRs across 10+ open source repos every single day. And this is only going to accelerate. everything is collapsing into a question of money.
14
25
1,449