Interested in the latest research on AI evaluations and heading to #ACL2026 or #ICML2026? ✈️ We went through the accepted papers and curated a selection worth checking out if you’re interested in #AI #evaluations, #benchmarking, evaluation #methodology, and #measurement. 🧵👇

Jul 2, 2026 · 6:07 PM UTC

2
9
41
3,072
RelevantRecentLikes
📄 #ACL 2026 highlights (1/2) • AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs • SCAN: Structured Capability Assessment and Navigation for LLMs • SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA
1
3
224
📄 #ACL2026 highlights (2/2) • HoWToBench: Holistic Evaluation for LLM’s Capability in Human-level Writing • TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
1
2
212
📄 #ICML2026 highlights (1/2) • Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks • Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models • Noise Tectonics: Measuring the Stability of AI Benchmark Ecosystems
1
139
📄 ICML 2026 (2/2) • Quantifying Biases in LLM-as-a-Judge Evaluations • STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems • AI Evaluations Should be Grounded on a Theory of Capability • Agent Evaluation Should Be Agentified for Openness, Standardization,…
1
2
143
🚀 Plus our own #EvalEval papers at #ICML2026 @icmlconf: • When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation • Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
1
1
164
Check out the full accepted paper lists: 📚 @aclmeeting 2026: 2026.aclweb.org/program/acce… 📚 @icmlconf 2026: icml.cc/virtual/2026/papers.… Did we miss any interesting papers on AI evaluations? Share your recommendations! ⬇️
1
132
bookmarking this so i can pretend i understand any of it when the small-cap software i own eventually claims to have "ai evaluation infrastructure"
12