Research
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
arXiv:2509.21267v3 Announce Type: replace Abstract: Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, re
arXiv:2509.21267v3 Announce Type: replace Abstract: Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of problem-solving strategy but should maintain the same verifiable answer. Whereas, for creative writing tasks, we often expect variation in key narrative components (e.g. plot, setting, etc.) beyond mere vocabulary diversity. Prior work on homogenization rarely conceptualizes diversity in a task-dependent way. We address this gap with four contributions: (1) a task taxonomy with distinct notions of functional diversity -- whether a user would perceive two responses as meaningfully different for a given task; (2) a small user study validating that the taxonomy aligns with human perception of functional diversity; (3) a task-dependent sampling technique that increases diversity only where homogenization is undesired; (4) evidence challenging the perceived diversity-quality trade-off, showing it may stem from mis-conceptualizing both diversity and quality in a task-agnostic way.
Related
- Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios
- Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
- Adaptive Multi-Expert Reasoning via Difficulty-Aware Routing and Uncertainty-Guided Aggregation
- CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
Source: arXiv cs.CL | 2026-04-23