Model Releases
FlowInOne - A new Multimodal image model . Released on Huggingface
FlowInOne is a vision-centric multimodal image generation framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean i...
FlowInOne is a vision-centric multimodal image generation framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean image-in, image-out pipeline governed by a single flow matching model. It unifies text-to-image generation, layout-guided editing, and visual instruction following under one coherent paradigm, eliminating cross-modal alignment bottlenecks and task-specific architectural branches. Released on Hugging Face in April 2026 (arXiv: 2604.06757), it achieves state-of-the-art performance across unified generation tasks, surpassing both open-source models and competitive commercial systems, and is accompanied by the VisPrompt-5M dataset of 5 million visual prompt pairs and the VP-Bench evaluation benchmark.
Related
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
- Visual prompting reimagined: The power of the Activation Prompts
- PLUME: Latent Reasoning Based Universal Multimodal Embedding
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
Source: model-releases