Model Releases
Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer
arXiv:2604.11687v1 Announce Type: new Abstract: AI-generated text has become common in academic and professional writing, prompting research into detection methods. Less studied is the reverse: system
arXiv:2604.11687v1 Announce Type: new Abstract: AI-generated text has become common in academic and professional writing, prompting research into detection methods. Less studied is the reverse: systematically rewriting AI-generated prose to read as genuinely human-authored. We build a parallel corpus of 25,140 paired AI-input and human-reference text chunks, identify 11 measurable stylistic markers separating the two registers, and fine-tune three models: BART-base, BART-large, and Mistral-7B-Instruct with QLoRA. BART-large achieves the highest reference similarity -- BERTScore F1 of 0.924, ROUGE-L of 0.566, and chrF++ of 55.92 -- with 17x fewer parameters than Mistral-7B. We show that Mistral-7B's higher marker shift score reflects overshoot rather than accuracy, and argue that shift accuracy is a meaningful blind spot in current style transfer evaluation.
Related
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
- Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network
- Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences
- Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry
- Many Ways to Be Fake: Benchmarking Fake News Detection Under Strategy-Driven AI Generation
Source: arXiv cs.CL | 2026-04-14