NTH

GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

AuthorsXiangtao Kong, Jixin Zhao, Lingchen Sun, Rongyuan Wu, Lei Zhang

June 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper uses multimodal foundation models to fabricate realistic ground-truth images from degraded photos, creating a 100K-pair dataset that helps restoration models work better on real-world images.

Key results

103707
training_pairs

GGT-100K contains 103,707 training pairs built from real-world low-quality images and generative high-quality targets.

500
test_pairs

The dataset includes a curated 500-pair test set.

1024
resolution

All candidate images are normalized to 1024 × 1024 resolution for dataset construction.

0.84
best_overall_score

Nano-Banana-2 with Gemini-based adaptive prompting achieved the best overall Avg. score in the MFM benchmark.

32.5%
human_preference

Nano-Banana-2 with Gemini-based adaptive prompting received the highest human preference in the user study.

3.12
NAFNet_PSNR_gain

Training NAFNet with GGT-100K improved test-set PSNR by 3.12 dB.

What the paper found

GGT-100K proposes a new data paradigm for real-world image restoration by using multimodal foundation models, rather than physical capture or synthetic degradations, to generate “generative ground truth” targets from low-quality images. The authors, from The Hong Kong Polytechnic University and OPPO Research Institute, systematically benchmarked nine MFMs, including OpenAI’s GPT-Image-1.5 and GPT-Image-2, Google’s Gemini-based prompting, Black Forest Labs’ FLUX.2-dev, and Nano-Banana-2, across fidelity, perceptual quality, VLM-based acceptance, and human preference. Nano-Banana-2 with Gemini-based adaptive prompting was selected as the best balance, achieving the top overall score of 0.84 and the highest human preference at 32.5%. Using this model, they built GGT-100K with 103,707 training pairs at 1024×1024 resolution and a 500-pair test set, spanning general mixed corruption, rain, haze, snow, low-light, and old photos. A three-stage quality-control pipeline—no-reference metric filtering, VLM-assisted refinement with Gemini-3.1-Pro, and manual verification—removed hallucinations and structural inconsistencies. Training on GGT-100K consistently improved generalization across CNN, Transformer, all-in-one, and generative restoration models; for example, NAFNet gained 3.12 dB PSNR and +26.2 percentage points VLM-R, while X-Restormer improved by 3.54 dB and +24.2 points. The largest benefits appeared in generative models such as FLUX-Controlnet and Qwen-Image-Edit, showing that foundation-model-generated supervision can expand restoration beyond the limits of existing paired datasets.

Original abstract

Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail to model real-world degradations, while real-world paired datasets are expensive and difficult to capture. As a result, IR models trained on these datasets show limited generalization in real-world scenarios. In this work, we propose Generative Ground Truth (GGT) by using generative multimodal foundation models (MFMs) to produce high-quality (HQ) targets from real-world low-quality (LQ) images. We first conduct a systematic evaluation of nine state-of-the-art MFMs, including Nano-Banana-2 and GPT-Image-2, on images of various scenes and degradation types. The results demonstrate that Nano-Banana-2 with VLM-based adaptive prompting shows the highest capability to synthesize perceptually realistic and content-faithful HQ targets, which can serve as the GGT for the LQ input. We then employ Nano-Banana-2 to build a GGT synthesis pipeline, which involves multi-stage quality control to ensure data reliability, and construct GGT-100K, an LQ-HQ paired dataset comprising 103,707 training pairs and covering diverse scenes and complex real-world degradations. A test set of 500 image pairs is also established. Extensive experiments show that GGT-100K consistently improves the real-world generalization of a wide range of IR models, with particularly strong benefits for finetuning generative models for IR tasks. Our results suggest that MFMs can serve as practical tools for restoration-oriented data generation, and GGT-100K is a useful resource to expand the generalization boundaries of real-world IR models.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis