Masked Generative Transformer Is What You Need for Image Editing

Journal: arXiv
Published Date:

Abstract

Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications to propagate into areas that should remain intact. We propose a fundamentally different approach by leveraging Masked Generative Transformers (MGTs), whose localized token-prediction paradigm naturally confines changes to intended regions. We present EditMGT, an MGT-based editing framework that is the first of its kind. Our approach employs multi-layer attention consolidation to aggregate cross-attention maps into precise edit localization signals, and region-hold sampling to explicitly prevent token flipping in non-target areas. To support training, we construct CrispEdit-2M, a 2M-sample high-resolution (>1024) editing dataset spanning seven categories. With only 960M parameters, EditMGT achieves state-of-the-art image similarity on multiple benchmarks while delivering 6x faster editing, demonstrating that MGTs offer a compelling alternative to diffusion-based editing.

Authors

  • Wei Chow; Linfeng Li; Xian Sun; Lingdong Kong; Zefeng Li; Qi Xu; Hang Song; Tian Ye; Xian Wang; Jinbin Bai; Shilin Xu; Xiangtai Li; Junting Pan; Shaoteng Liu; Ran Zhou; Tianshu Yang; Songhua Liu