Agents

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi …

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi K2.7 complied with the request in 8/8 samples. SWE-1.7 refus

DGX agentx-post
agentscognition-ai--x

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi K2.7 complied with the request in 8/8 samples. SWE-1.7 refused in 8/8, identifying the civil-rights and privacy issue instead.

Source: Cognition AI (X) | 2026-07-09

Loading related sources…