Industry

Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA…

Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA modifies the RoPE encodings during the diffusion process to

DGX agentx-post
industryemad-mostaque--x

Image diffusion models like Flux natively output at 1k resolution, but what if we want to generate much higher resolution images (6k+)? SEGA modifies the RoPE encodings during the diffusion process to generate high-resolution images---no fine-tuning required! 📢🚀 Happy to share our recent work. SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers 📄Paper: https://arxiv.org/abs/2605.22668 🌐Webpage: https://rajabi2001.github.io/sega/ 🤗HF: https://huggingface.co/papers/2605.22668

Source: Emad Mostaque (X) | 2026-05-23

Loading related sources…