@sophronresearchi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
We develop evaluations to empower human understanding and decision-making in a world of advanced AI | Developers of the Pander Score
New York
Joined August 2026
- Tweets12
- Following2
- Followers261
- Likes9
News: Astra gets a near-zero result on the Pander Score.
New scores are in for GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash. Gemini is virtually unchanged from 3.7 Flash, while Fable 5.1 improves on instructions relative to 5. These changes are substantially smaller than Astra relative to Sol.
Astra's biggest jump is on the instructional Pander Score: whether a model will call out dubious presuppositions in instructions when appropriate. No other model is close to a 0 score on instructional prompts.
This means that the Pander Score detects no substantial pandering in Astra.
Note: This is no guarantee that future GPT models will score similarly. It also does not mean that Astra does not pander, only that our current methods do not detect it. We are currently working on a multi-turn evaluation that will provide higher signal.
Here are the Pander Score results on Anthropic's most recent models, compared with Opus and Sonnet 4.6. Some takeaways:
Claude models mostly outperform others, with competition from Meta's Muse Spark 1.1, Moonshot AI's Kimi K3, and OpenAI's GPT-5.6 Sol.
More capable models need not be less sycophantic. Sonnet 5 panders more than Sonnet 4.6.
Opus and Sonnet 4.6 are slightly contrarian in conversation, meaning that they are less likely to agree with what the user seems to believe.
Opus 5 outperforms all other models on instruction prompts. This means that it is less likely to uncritically accept factual assumptions as background for a task.
Full results can be found at sophronresearch.org/pander
Sophron Research has officially launched. We develop evaluations to empower human understanding and decision-making in a world of advanced AI.
Today we are launching Sophron Research (@sophronresearch): an independent research nonprofit developing evaluations for AI models. Our mission is to empower human understanding and decision-making in a world of advanced AI.
As AI systems become more capable, we will rely on them to inform and execute decisions in our lives and integrate them into our institutions.
This can go badly. In one future world, models interact with us in ways that undermine our judgment and autonomy. They might convince us to believe something that would benefit their developer, or keep us engaged by sycophantically telling us what we want to hear. Or they might perform most functions in society in ways that are completely opaque to humans, leaving us disempowered.
But there is another world we can aim for. In this world, models empower people by providing reliable and understandable advice, and accurately explain their actions in verifiable ways. As a consequence, we continue to scale our own understanding and capabilities with those of our models.
Sophron exists to help steer us towards the second world. We do this by developing evaluations to assess whether AI models support or undermine sound judgment. Drawing on formal philosophy, statistics, and cognitive science, we identify qualities like accuracy and honesty and turn them into concrete scores that developers and policymakers can act on. Our first evaluation, the Pander Score, measures whether models adapt their views to what users already seem to believe.
We pursue this as a nonprofit third-party initiative. AI developers will not always have incentives to build products that improve our autonomy. Our aim is to provide accountability through independent tests, making transparent to developers, policymakers, and consumers how the models behave. We publish our methods, data, and results for free online.
New evaluations, results, and opportunities will be shared on @sophronresearch, so follow us there to stay posted. Sophron is founded by @PReaulx and @alejbo.
The name comes from the Greek word 'sophron', meaning 'of sound mind'.
Sophron Research retweeted
How does an AI model's expressed belief depend on the user's expressed belief?
Great work defining and evaluating epistemic sycophancy by @sophronresearch. I think this is a consequential AI behavior that has been challenging for researchers to operationalize well.
Replying to @PReaulx
Today we're launching the Pander Score 🐼: a public and continuously updated sycophancy leaderboard, measuring how much AIs shift their views to agree with users.
High score = the AI mirrors your views. 0 score = the AI is independent.
The differences between current flagship models are large. Claude Fable 5 performs the best, basically ignoring the user's view entirely, while GLM-5.2 notably adapts its responses to agree with users.
Other models fall in between, with Muse Spark 1.1, GPT 5.6 Sol, and Kimi K3 doing better than Grok 4.6, Gemini 3.7 Flash, and Inkling.
Apply to work with Sophron through the @SPARexec mentorship program. Deadline Aug 21.
Do you or someone you know want to do research on sycophancy and related epistemic evaluations? @sophronresearch is participating in @SPARexec mentorship program. Consider applying to our stream; deadline Aug 21. Link below.
Sophron Research retweeted
Ask an AI a question and it might agree with what you already seem to believe, whether or not you're right. If so, it panders to you.
Here is Gemini 3.5 Flash asked about Reiki energy healing, giving opposite answers when asked by a skeptic vs. a believer. 🧵