Industry
Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA…
Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA modifies the RoPE encodings during the diffusion process to
Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA modifies the RoPE encodings during the diffusion process to generate high-resolution images---no fine-tuning required! 📢🚀 Happy to share our recent work. SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers 📄Paper: https://arxiv.org/abs/2605.22668 🌐Webpage: https://rajabi2001.github.io/sega/ 🤗HF: https://huggingface.co/papers/2605.22668
Source: Emad Mostaque (X) | 2026-05-23