Model Releases

On FrontierCode 1.1 Extended, our benchmark for real-world engineering tasks that grades mergeability and quality, Kimi K3 scores 58.2% with…

On FrontierCode 1.1 Extended, our benchmark for real-world engineering tasks that grades mergeability and quality, Kimi K3 scores 58.2% with a 63.6% pass rate. Within Devin, it excels on reproducing b

DGX agentx-post
model-releasescognition-ai--x

On FrontierCode 1.1 Extended, our benchmark for real-world engineering tasks that grades mergeability and quality, Kimi K3 scores 58.2% with a 63.6% pass rate. Within Devin, it excels on reproducing bugs and managing its environment effectively. https://devin.ai/blog/kimi-k3

Source: Cognition AI (X) | 2026-07-27

Loading related sources…