In this video, we'll look at the evolution of neural networks for generating and editing images. We'll start with early models that created "multi-armed polypods" and move on to modern solutions like FluxContext, which can subtly modify images based on a text request, although not without drawbacks, such as facial degradation.
I'll talk about adding new popular models — NanoBanana from Google (based on Gemini) and the very economical QWEN Image Edit from Alibaba. Each has its own nuances: NanoBanana can ignore specified dimensions, and QWEN can't generate from scratch and requires the source.
The main idea of the video is to stop opposing AI and Photoshop, and combine them. I will show the concept of a future AI editor, where the language model (LLM) acts as a "brain" that analyzes the user's request and selects the right diffusion model for a specific task: be it quality improvement (upscaling), background removal, or drawing from a sketch.
You will see my attempt to implement this by embedding AI directly into the browser version of Photoshop (Photopea fork), and why I had to abandon this idea due to licensing restrictions. As an alternative, I present a solution based on the MiniPint editor with an open license, as well as a fork of the Excalidro vector editor with a built-in AI chat and an infinite canvas for free creativity. This is my view on the near future of professional graphics work.
Useful links:
Dewiar service: https://dewiar.com
NeuroTouch image generator: https://dewiar.com/neuro_touch/
Join our Telegram channel: https://t.me/dewiarx
Project support: Help improve the stability of the system - we are raising funds for a generator for uninterruptible power supply of servers:
https://donate.stream/donate_dewiar
Each donation goes to the development of infrastructure so that your automations work without failures 24/7.