researching AI’s societal impacts @TransluceAI 🔍 // 💭 AI (evals, governance, policy)
- Tweets478
- Following769
- Followers806
- Likes6.5K
ALT Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.