Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
DGX agentarXiv:2607.21988v1 Announce Type: new Abstract: Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enab