Editing Text in Images with AI. Research Review for Scene Text Editing… | by Julia Turc | Towards Data Science
If you ever tried to change the text in an image, you know it’s not trivial. Preserving the background, textures, and shadows takes a Photoshop license and hard-earned designer skills. In the video below, a Photoshop expert takes 13 minutes to fix a few misspelled characters in a poster that is not even stylistically complex. The good news is — in our relentless pursuit of AGI, humanity is also building AI models that are actually useful in real life. Like the ones that allow us to edit text in images with minimal effort. The task of automatically updating the text in an image is formally known as Scene Text Editing (STE). This article describes how STE model architectures have evolved over time and the capabilities they have unlocked. We will also talk about their limitations and the work that remains to be done. Prior familiarity with GANs and Diffusion models will be helpful, but not strictly necessary. Disclaimer: I am the cofounder of Storia AI, building an AI copilot for visual e
Editing Text in Images with AI | Towards Data Science Artificial Intelligence Editing Text in Images with AI Research Review for Scene Text Editing: STEFANN, SRNet, TextDiffuser, AnyText and more. Julia Turc Feb 18, 2024 17 min read Share If you ever tried to change the text in an image, you know it’s not trivial. Preserving the background, textures, and shadows takes a Photoshop license and hard-earned designer skills. In the video below, a Photoshop expert takes 13 minutes to fix a few misspelled characters in a poster that is not even stylistically complex. The good news is – in
Explore this link on the map →saved by
related reading
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversiontextual-inversion.github.io
- The Illustrated Stable Diffusion – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Introducing CM3leon, a more efficient, state-of-the-art generative model for text and imagesai.meta.com
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Unified Multimodal Models as Auto-Encodersarxiv.org
- ⭐️ Diffusion Modelsandrewkchan.dev
- Nano Banana can be prompt engineered for extremely nuanced AI image generation | Max Woolf's Blogminimaxir.com
- Replicate - Run AI with an APIreplicate.com
- How does Stable Diffusion work?stable-diffusion-art.com
- Turning off lights with model editing — AI Alignment Forumalignmentforum.org
- Large Language Diffusion Modelsarxiv.org
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org