Safety

Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

arXiv:2608.27280v1 Announce Type: new Abstract: Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form stor

DGX agentpaper
safetyarxiv-cs-cv

arXiv:2608.27280v1 Announce Type: new Abstract: Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introduced and is later referred to by a type-level mention or pronoun. Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain. We propose extbf{Sidecar}, a plug-and-play semantic augmentation module that preserves entity-level information from the initial description and injects the missing semantics into later prompt embeddings. Sidecar requires no additional training and does not modify the architecture of the base diffusion model. Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead.

Related

Source: arXiv cs.CV | 2026-08-28

Loading related sources…