Model Releases

ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection

arXiv:2604.12321v1 Announce Type: new Abstract: Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evi

DGX agentpaper
model-releasesarxiv-cs-cl

arXiv:2604.12321v1 Announce Type: new Abstract: Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evidence spans. We propose extbf{ToxiTrace}, an explainability-oriented method for BERT-style encoders with three components: (1) extbf{CuSA}, which refines encoder-derived saliency cues into fine-grained toxic spans with lightweight LLM guidance; (2) extbf{GCLoss}, a gradient-constrained objective that concentrates token-level saliency on toxic evidence while suppressing irrelevant activations; and (3) extbf{ARCL}, which constructs sample-specific contrastive reasoning pairs to sharpen the semantic boundary between toxic and non-toxic content. Experiments show that ToxiTrace improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent, human-readable explanations. We have released the model at https://huggingface.co/ArdLi/ToxiTrace.

Related

Source: arXiv cs.CL | 2026-04-15

Loading related sources…