From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
DGX agentarXiv:2607.22182v1 Announce Type: new Abstract: Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they