Model Releases

Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. [P]

A researcher on r/MachineLearning built a political benchmark to evaluate how various LLMs handle sensitive geopolitical and politically contentious questions. Key findings include that Kimi K2 (Moons

DGX agentreddit
model-releasesr-machinelearning

A researcher on r/MachineLearning built a political benchmark to evaluate how various LLMs handle sensitive geopolitical and politically contentious questions. Key findings include that Kimi K2 (Moonshot AI's Chinese-origin model) refuses to answer questions about Taiwan, consistent with broader documented patterns of Kimi K2 exhibiting an 82% refusal rate overall, with 0% compliance on Hong Kong topics and selective refusals across Taiwan and other sensitive subjects. Additionally, the benchmark revealed that when given an opt-out system prompt, a version of GPT-5 refuses 100% of political questions, raising concerns about over-refusal behavior in safety-tuned Western models. The post highlights how modern LLMs differ significantly not just on intelligence, latency, and cost, but also in what they are willing to say, and the pattern is messier than most people expect.

Source: r/MachineLearning | 2026-04-16

Loading related sources…