Model Releases

Quoting Anthropic

We used an automatic classifier which judged sycophancy by looking at whether Claude showed a willingness to push back, maintain positions when challenged, give praise proportional to the merit of ide

DGX agentarticle
model-releasessimon-willison

We used an automatic classifier which judged sycophancy by looking at whether Claude showed a willingness to push back, maintain positions when challenged, give praise proportional to the merit of ideas, and speak frankly regardless of what a person wants to hear. Most of the time in these situations, Claude expressed no sycophancy—only 9% of conversations included sycophantic behavior (Figure 2). But two domains were exceptions: we saw sycophantic behavior in 38% of conversations focused on spirituality, and 25% of conversations on relationships. — Anthropic, How people ask Claude for personal guidance Tags: ai-ethics, anthropic, claude, ai-personality, generative-ai, ai, llms, sycophancy

Source: Simon Willison | 2026-05-03

Loading related sources…