Local Ai
Why isn't JoyAI Image Edit getting any love?
A Reddit thread on r/StableDiffusion questioning why JoyAI Image Edit — a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing —
A Reddit thread on r/StableDiffusion questioning why JoyAI Image Edit — a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing — hasn't received more community attention. The model combines an 8B Multimodal Large Language Model (MLLM) with a 16B Multimodal Diffusion Transformer (MMDiT) , though its full bf16 transformer weighs 32.5 GB — too large for consumer GPUs like the RTX 4090 — with an FP8 quantized version reducing that to ~16 GB , which likely contributes to limited mainstream adoption despite its capabilities.
Related
- JoyAI-Image-Edit now has ComfyUI support
- how much of vram i need for joy-image-edit
- Spatial Edit (Apache 2.0)
- Forget about VAEs? SenseNova's NEO-unify achieves 31.5 PSNR without an encoder – Native Image Gen is coming.
Source: r/StableDiffusion | 2026-04-15