Model Releases
Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy
arXiv:2608.19006v1 Announce Type: new Abstract: Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and
arXiv:2608.19006v1 Announce Type: new Abstract: Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech detection (HSD) systems is paramount, we argue that this must also be balanced with the individual right to privacy. Exploring the intersection of HSD and privacy, we demonstrate that HSD systems might unintentionally achieve performance at the cost of encoding authorship, posing a threat to privacy. Building on these findings, we establish the notion of a privacy-HSD trade-off, which demands a careful balance. We benchmark a series of text privatization methods, as well as our newly proposed domain-specific AgnoSpeech technique, showing that balancing privacy and HSD is difficult but feasible. The findings make a strong case for more research on the trade-offs between privacy and HSD, both of which have tangible implications for the safeguarding of online participation.
Related
- Comparison of Modern Multilingual Text Embedding Techniques for Hate Speech Detection Task
- From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation
- A Comparative Study of PyCaret AutoML and CNN-BiLSTM for Binary Hate Speech Detection in Indonesian Twitter
Source: arXiv cs.CL | 2026-08-20