FashionComposer: 構成的ファッション画像生成

要旨

ファッション画像生成のための構成的なFashionComposerを提案します。従来の手法とは異なり、FashionComposerは非常に柔軟です。テキストプロンプト、パラメトリックな人間モデル、衣服画像、および顔画像といったマルチモーダルな入力を受け入れ、人物の外見、ポーズ、体型を個人化し、1度の処理で複数の衣服を割り当てることができます。これを実現するために、まず多様な入力モダリティを処理できる汎用フレームワークを開発します。モデルの堅牢な構成能力を向上させるために、スケーリングされたトレーニングデータを構築します。複数のリファレンス画像（衣服や顔）をシームレスに取り込むために、これらのリファレンスを「アセットライブラリ」として1つの画像に整理し、外見特徴を抽出するためにリファレンスUNetを使用します。生成された結果の正しいピクセルに外見特徴を注入するために、サブジェクトバインディングアテンションを提案します。これにより、異なる「アセット」からの外見特徴を対応するテキスト特徴と結びつけます。この方法により、モデルは各アセットの意味に基づいて理解し、任意の数や種類のリファレンス画像をサポートします。包括的なソリューションとして、FashionComposerは人物アルバム生成、多様なバーチャル試着タスクなど、他の多くのアプリケーションもサポートしています。

English

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.

FashionComposer: 構成的ファッション画像生成

FashionComposer: Compositional Fashion Image Generation

要旨

Summary

Support

Support