Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
DGX agentarXiv:2602.10388v3 Announce Type: replace-cross Abstract: The diversity of post-training data is critical for effective downstream performance in large language models (LLMs). Many existing approaches