PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
DGX agentarXiv:2605.17028v1 Announce Type: cross Abstract: Large language models (LLMs) hallucinate with confidence: their outputs can be fluent, authoritative, and simply wrong. In medical, legal, and scienti