Improved and Accelerated Text-to-Image Generation with Collect, Reflect, and Refine.

Shao, Shitong; Zhou, Zikai; Xie, Dian; Fang, Yuetong; Ye, Tian; Bai, Lichen; Han, Bo; Xie, Zeke · IEEE Trans Pattern Anal Mach Intell · 2026

basic_science · Level V

Where this comes from

Abstract

Recently, enhancing the generative capability of text-to-image (T2I) models has become a promising direction in both academia and industry. Prior studies often focused on either improving generative quality or reducing inference latency, but typically failed to improve both quality and speed simultaneously. Moreover, existing inference-enhancement methods do not achieve significant improvements simultaneously across both diffusion models (DMs) and autoregressive models (ARMs). In this paper, we introduce a general tuning-based inference-enhancement framework, named ${\bm {CoRe}^{2}}$, which is the first to simultaneously achieve significant generative quality and reduced inference overhead across DMs and ARMs, to the best of our knowledge. ${\bm {CoRe}^{2}}$ comprises three stages: ${\bm {Co}llect}$, ${\bm {Re}flect}$, and ${\bm {Re}fine}$. During the Collect stage, classifier-free guidance (CFG) trajectories are collected and subsequently used in the Reflect stage to train a weak model capable of reflecting the "easy-to-learn" content. Finally, during the Refine stage, CoRe$^{2}$ can utilize the trained weak model to achieve speedup and performance gain in inference. Specifically, in the early sampling steps, CoRe$^{2}$ employs weak-to-strong guidance to refine the "difficult-to-learn" and realistic content, thereby improving generative quality. In the later sampling steps, CoRe$^{2}$ can use the weak model to generate "easy-to-learn" content instead of CFG, dramatically reducing inference time. Experimental outcomes substantiates CoRe$^{2}$ achieve significant performance improvements on HPD v2, Pick-of-Pic, Drawbench, GenEval, and T2I-Compbench across SDXL, SD3.5, FLUX and LlamaGen. Notably, for SD3.5, CoRe$^{2}$ can be seamlessly integrated with the state-of-the-art inference-enhancement algorithm Z-Sampling, outperforming it with winning rates of 73% and 69% on PickScore and AES even with less time.