FreeMD: Training-free multi-domain text-to-image generation with any control.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

Diffusion models have promoted the development of controllable text-to-image generation with structure control. However, there still remain two limitations: (1) existing methods coupling text prompt and structure control inevitably lead to pixel-level structure control dominates the generation process, thus misalignment with text prompt; (2) they suffer poor structure consistency due to the fact that these methods typically focus only on spatial domain features, while neglecting the essential frequency domain representations including texture detail. To alleviate the above issues, we propose FreeMD, a novel training-free multi-domain text-to-image generation method to realize better semantic alignment with text prompt and achieve excellent structure consistency with structure control simultaneously. Specifically, we design two independent guidance branches to decouple text prompt and structure control: appearance guidance branch and structure guidance branch. The former utilizes principal component supervision for elaborate appearance representations to transfer appearance in text to the generated image. Such a skillful paradigm is capable of facilitating semantic alignment with text prompt in generation process. Collaboratively, the latter designs a multi-domain guidance strategy combining spatial domain and frequency domain by comprehensive supervision for structure representations, thus improving structure consistency with control. Thanks to the decoupled architecture and multi-domain guidance strategy, FreeMD accurately aligns with text prompt as well as achieves structurally coherent with control signals. Moreover, FreeMD can be plug-and-play in various pre-trained generative models to accomplish common downstream tasks. Extensive experiments demonstrate FreeMD outperforms the existing methods in controllability and generation quality.

Authors

Keywords

No keywords available for this article.