Interested in the latest research on AI evaluations and heading to #ACL2026 or #ICML2026? ✈️
We went through the accepted papers and curated a selection worth checking out if you’re interested in #AI #evaluations, #benchmarking, evaluation #methodology, and #measurement. 🧵👇
Jul 2, 2026 · 6:07 PM UTC
📄 #ACL 2026 highlights (1/2)
• AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
• SCAN: Structured Capability Assessment and Navigation for LLMs
• SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA
📄 #ACL2026 highlights (2/2)
• HoWToBench: Holistic Evaluation for LLM’s Capability in Human-level Writing
• TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
📄 #ICML2026 highlights (1/2)
• Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
• Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models
• Noise Tectonics: Measuring the Stability of AI Benchmark Ecosystems
📄 ICML 2026 (2/2)
• Quantifying Biases in LLM-as-a-Judge Evaluations
• STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
• AI Evaluations Should be Grounded on a Theory of Capability
• Agent Evaluation Should Be Agentified for Openness, Standardization,…
Check out the full accepted paper lists:
📚 @aclmeeting 2026: 2026.aclweb.org/program/acce…
📚 @icmlconf 2026: icml.cc/virtual/2026/papers.…
Did we miss any interesting papers on AI evaluations? Share your recommendations! ⬇️