Model Releases
How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode
Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a
Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a benchmarking issue? EDIT: The contribution of this benchmark to the Intelligence index of artificialanalysis.ai: Full Intelligence Index v4.1 weights: GDPval-AA v2: 20% Terminal-Bench 2.1: 16% τ³-Bench Banking: 14% Humanity's Last Exam: 12% AA-Omniscience Accuracy: 8% SciCode: 8% GPQA: 6% AA-LCR: 6% CritPt: 6% AA-Omniscience Non-Hallucination: 4% Source: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1 submitted by /u/Informal-Trouble2183 [link] [comments]
Related
- Current smallest usable coding model
- DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark
- Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)
- SWE-rebench Multilingual Update (Go, Java, Python, Rust, TS). Evaluated: GLM-5.2, DeepSeek-V4 Pro, Qwen3.6-27B and others
Source: r/LocalLLaMA | 2026-08-06