Model Releases

resharing this note, find it helpful given all the great open evals work + teams building vertical agents Evals are a proxy for the behavior…

resharing this note, find it helpful given all the great open evals work + teams building vertical agents Evals are a proxy for the behavior we want our agent to exhibit in production Model+Harness pu

DGX agentx-post
model-releasesharrison-chase--x

resharing this note, find it helpful given all the great open evals work + teams building vertical agents Evals are a proxy for the behavior we want our agent to exhibit in production Model+Harness public benchmark scores only really matter if their task distribution accurately reflects the work your agent needs to do Look at the data there's big alpha in reading the actual eval tasks + output traces and thinking about if these map onto the product experience you want for your customers if the capabilities measured in the bench diverge from your use-case and required model abilities, then the signal from the score is not a great proxy for your agent's perf/behavior Model-Harness-Task fit means that each of these axes need to be aligned to create a good agent experience. Harnesses need to provide the right tools to do a task in the right format the model natively expects from post-training Models need to have good instructions and be beyond the intelligence threshold for the task A good way to do this is: - Building bespoke Evals that measure your use case specifically instead of generic Evals - Dogfooding & Human review of traces - Sourcing and measuring signal from production traces which capture the true distribution of how people are using your agent lots of alpha to extract in vertical agents by carefully co-designing and creating feedback/update loops between model choice, harness design, and task evals reach out if we can help you do that! :)

Source: Harrison Chase (X) | 2026-04-30

Loading related sources…