NTH

RefGC-SR$^2$: Reference-guided Generated Content Super-Resolution and Refinement

AuthorsJeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh

June 22, 2026 3 min read
Watch on YouTube
The one-line take

This work proposes a new way to restore and sharpen AI-generated images using the original reference image, recovering lost fine details while fixing generation artifacts in one step.

Key results

40K
RefGC-SR2 training triplets

Size of the RefGC-SR2 Dataset used for supervised training

200
RefGC-SR2 benchmark samples

Evaluation split size for the RefGC-SR2 Benchmark

0.8696
CLIP-I

Best overall score on the RefGC-SR2 benchmark

0.7474
DINO

Best overall score on the RefGC-SR2 benchmark

17.5148
PSNR

Best overall score on the RefGC-SR2 benchmark

0.6335
SSIM

Best overall score on the RefGC-SR2 benchmark

What the paper found

RefGC-SR2, from Chung-Ang University and Adobe Research, defines a new post-processing task for reference-guided generation: given a low-resolution generated image and the original high-resolution reference image, the model must simultaneously recover lost detail, upscale the output, and remove generative artifacts such as identity distortion, texture loss, and detail inconsistency. To make this possible, the authors build the first real-world LRGI-HRRI-HRGT triplet dataset, RefGC-SR2 Dataset, with 40K training triplets and a 200-sample benchmark, and synthesize the low-quality anchors using DipRefGC, a FLUX-based diptych-conditioned generator with dual ControlNets that enforces pose consistency while preserving reference appearance. The proposed RefGC-SR2 model is built on FLUX-Kontext and injects frequency-adaptive Mixture-of-LoRA Experts across all DiT blocks, routing low-frequency structure early and high-frequency detail late, while a frequency-based loss aligns low-frequency latent content to HRGT and high-frequency statistics to HRRI. On the RefGC-SR2 benchmark, it achieves the best overall scores, including CLIP-I 0.8696, DINO 0.7474, PSNR 17.5148, SSIM 0.6335, and LPIPS 0.2746, and in user studies it is ranked first by 82–83 percent of participants across refinement, detail, and overall quality, showing that the method is not just sharper but also more faithful to the reference and more robust on real compositing and customization outputs from models such as Gemini 2.5 Flash Image, GPT-Image 1.5, and Qwen-Image-Edit.

Original abstract

Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object-centric high-resolution reference image (HRRI) provided by users is downsampled to a fixed low-resolution (LR) before being fed into the model, so the fine-grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference-guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference-guided super-resolution (RefSR) methods recover resolution but assume natural-image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference-guided generated content super-resolution-refinement (RefGC-SR$^2$), where the original HRRI is reused at the post-processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real-world triplet data generation pipeline for this RefGC-SR$^2$ task, training a diptych-conditioned generator to synthesize paired low-quality anchors that public pretrained models cannot provide. We further present a frequency-aware diffusion transformer model for RefGC-SR$^2$ that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC-SR$^2$ model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high-resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis