RefGC-SR$^2$: Reference-guided Generated Content Super-Resolution and Refinement
AuthorsJeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh
Resources
This work proposes a new way to restore and sharpen AI-generated images using the original reference image, recovering lost fine details while fixing generation artifacts in one step.
Key results
Size of the RefGC-SR2 Dataset used for supervised training
Evaluation split size for the RefGC-SR2 Benchmark
Best overall score on the RefGC-SR2 benchmark
Best overall score on the RefGC-SR2 benchmark
Best overall score on the RefGC-SR2 benchmark
Best overall score on the RefGC-SR2 benchmark
What the paper found
RefGC-SR2, from Chung-Ang University and Adobe Research, defines a new post-processing task for reference-guided generation: given a low-resolution generated image and the original high-resolution reference image, the model must simultaneously recover lost detail, upscale the output, and remove generative artifacts such as identity distortion, texture loss, and detail inconsistency. To make this possible, the authors build the first real-world LRGI-HRRI-HRGT triplet dataset, RefGC-SR2 Dataset, with 40K training triplets and a 200-sample benchmark, and synthesize the low-quality anchors using DipRefGC, a FLUX-based diptych-conditioned generator with dual ControlNets that enforces pose consistency while preserving reference appearance. The proposed RefGC-SR2 model is built on FLUX-Kontext and injects frequency-adaptive Mixture-of-LoRA Experts across all DiT blocks, routing low-frequency structure early and high-frequency detail late, while a frequency-based loss aligns low-frequency latent content to HRGT and high-frequency statistics to HRRI. On the RefGC-SR2 benchmark, it achieves the best overall scores, including CLIP-I 0.8696, DINO 0.7474, PSNR 17.5148, SSIM 0.6335, and LPIPS 0.2746, and in user studies it is ranked first by 82–83 percent of participants across refinement, detail, and overall quality, showing that the method is not just sharper but also more faithful to the reference and more robust on real compositing and customization outputs from models such as Gemini 2.5 Flash Image, GPT-Image 1.5, and Qwen-Image-Edit.
Original abstract
Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object-centric high-resolution reference image (HRRI) provided by users is downsampled to a fixed low-resolution (LR) before being fed into the model, so the fine-grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference-guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference-guided super-resolution (RefSR) methods recover resolution but assume natural-image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference-guided generated content super-resolution-refinement (RefGC-SR$^2$), where the original HRRI is reused at the post-processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real-world triplet data generation pipeline for this RefGC-SR$^2$ task, training a diptych-conditioned generator to synthesize paired low-quality anchors that public pretrained models cannot provide. We further present a frequency-aware diffusion transformer model for RefGC-SR$^2$ that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC-SR$^2$ model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high-resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.