Relation Rectification in Diffusion Model
Despite their exceptional generative abilities, large text-to-image diffusion models, much like skilled but careless artists, often struggle with accurately depicting visual relationships between objects. This issue, as we uncover through careful analysis, arises from a misaligned text encoder that struggles to interpret specific relationships and differentiate the logical order of associated objects. To resolve this, we introduce a novel task termed Relation Rectification, aiming to refine the model to accurately represent a given relationship it initially fails to generate. To address this, we propose an innovative solution utilizing a Heterogeneous Graph Convolutional Network (HGCN). It models the directional relationships between relation terms and corresponding objects within the input prompts. Specifically, we optimize the HGCN on a pair of prompts with identical relational words but reversed object orders, supplemented by a few reference images. The lightweight HGCN adjusts the
Relation Rectification in Diffusion Model National University of Singapore International Conference on Computer Vision and Pattern Recognition (CVPR), 2024 *Corresponding Author. (a) Our approach enables diffusion model to successfully generate images with the correct directional relation in response to the textual prompt, which they originally failed. (b) Our method can synthesize relation of diverse and unseen objects in zero-shot manner. Abstract Despite their exceptional generative abilities, large text-to-image diffusion models, much like skilled but careless artists, often…
saved by
related reading
- Universal Guidance for Diffusion Modelsarxiv.org
- openaccess.thecvf.com/content/ICCV2023/papers/Li_Your_Diffusion_Model_is_Secretly_a_Zero-Shot_Classifier_ICCV_2023_paper.pdfopenaccess.thecvf.com
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondencesd-complements-dino.github.io
- One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusionarxiv.org
- What are Diffusion Models?lilianweng.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- ⭐️ Diffusion Modelsandrewkchan.dev
- VectorFusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Modelsajayj.com
- Fusing Diffusion Paths for Controlled Image Generationmultidiffusion.github.io
- A Dive into Text-to-Video Modelshuggingface.co