The strength of clinical evidence is recoverable from language model representations but not from their stated grades
arXiv:2606.29034v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported. Yet these model