Local Ai

Why isn't JoyAI Image Edit getting any love?

A Reddit thread on r/StableDiffusion questioning why JoyAI Image Edit — a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing —

DGX agentreddit
local-air-stablediffusion

A Reddit thread on r/StableDiffusion questioning why JoyAI Image Edit — a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing — hasn't received more community attention. The model combines an 8B Multimodal Large Language Model (MLLM) with a 16B Multimodal Diffusion Transformer (MMDiT) , though its full bf16 transformer weighs 32.5 GB — too large for consumer GPUs like the RTX 4090 — with an FP8 quantized version reducing that to ~16 GB , which likely contributes to limited mainstream adoption despite its capabilities.

Related

Source: r/StableDiffusion | 2026-04-15

Loading related sources…