Safety
EvoTale: Continual Character Customization for Expanding Story Worlds
arXiv:2603.16285v2 Announce Type: replace Abstract: Character-centric story visualization aims to synthesize coherent image sequences that depict narrative events and interactions while preserving rec
arXiv:2603.16285v2 Announce Type: replace Abstract: Character-centric story visualization aims to synthesize coherent image sequences that depict narrative events and interactions while preserving recurring character identities. In expanding story worlds, new user-specified characters must be continually incorporated despite varying customization difficulty and identity conflicts in multi-character scenes, without disrupting previously learned identities. In this paper, we propose EvoTale, a continual character customization framework for expanding stylized story worlds. We first introduce an All-in-One-World Character Integrator, which accumulates character-specific residual components within a unified LoRA branch using sparsely overlapping subspaces spanned by a shared orthonormal basis, while freezing previously learned components to limit cross-character coupling. We then develop a Character Quality Gate that uses rubric-guided MLLM feedback as a bounded controller to adapt the optimization budget based on the assessed customization quality. Finally, we propose Character-Aware Region-Focus Sampling, which combines bounding-box-guided regional denoising with identity-aware global denoising to preserve character identities within their designated regions while maintaining global narrative coherence. Experimental results show that EvoTale achieves a favorable balance across character fidelity, continual identity retention, multi-character generation quality, and story-text alignment compared with representative story visualization and customization methods.
Related
- Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
- Customizing Video Portraits via Identity-ActionDecoupling
- Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
Source: arXiv cs.CV | 2026-08-14