RefRetouchPaper ↗

ECCV 2026

RefRetouch
Personalized Image Retouching
without Test-time Fine-tuning

Temesgen Muruts Weldengus1Binnan Liu1Fei Kou2Youwei Lyu2Jinwei Chen2Qingnan Fan2†Changqing Zou1,3†
1 Zhejiang University2 vivo BlueImage Lab3 Zhejiang Lab† Corresponding authors
Single-reference, multi-reference, and photorealistic style-transfer results from RefRetouch
RefRetouch personalizes retouching from one or multiple references without test-time fine-tuning, and also supports photorealistic style transfer out of the box.

Motivation

Image retouching is subjective.

Users have different retouching preferences. Prior methods such as AdaInt [1], RSFNet [2], and DiffRetouch [3] learn a fixed style from one or several experts' paired data. A new user or style requires additional data and retraining.

Personalized retouching should instead adapt to a user on the go from a few exemplars that show the user's preference: tuning-free image retouching personalization.

1

Understand the preference

Encode the retouching difference in the reference pairs as a retouching latent.

2

Apply it adaptively

Select and combine the relevant preference exemplars for each new image.

Adobe Lightroom and Darktable mainly apply color and tone transformations. The main difference between input and retouched images is therefore in color and tone, including on MIT-Adobe FiveK [4] and PPR10K [5].

How can we encode this difference for tuning-free personalization?

Method · 01

Asymmetric auto-encoder

A LoRA-adapted Siamese SigLIPv2 encoder [8] processes the original and retouched images and represents their color and tone transformation as a retouching latent.

A lightweight conditional MLP maps each original RGB value to its retouched RGB value. Because this decoder operates only in color space, the latent must carry the retouching transformation rather than image content.

Asymmetric auto-encoder and retrieval-augmented retouching architecture
The asymmetric auto-encoder learns a retouching latent; RAR retrieves and aggregates relevant latents when multiple references are available.

Results · Retouching latent

Does the latent encode the user's edit?

01

Reconstruction

RefRetouch reconstructs the paired target more faithfully than MSM [6] on all three datasets.

DatasetMSMMSM retrainedRefRetouch
VCIRB27.28 / .900 / .09929.61 / .923 / .06932.42 / .974 / .035
MIT-FiveK28.94 / .851 / .13726.32 / .803 / .19932.95 / .972 / .028
PPR10K-A29.28 / .943 / .07529.71 / .936 / .08634.14 / .984 / .019

PSNR ↑ / SSIM ↑ / LPIPS ↓

02

Single-reference personalization

One reference latent transfers the retouching style to unseen images. On VCIRB, RefRetouch reaches 29.13 dB PSNR, compared with 24.35 dB for retrained MSM and 21.00 dB for StarEnhancer [7].

Single-reference retouching results with references, queries, outputs, and targets
03

Design ablation

Task adaptation, a color-space MLP, and additive conditioning give the strongest reconstruction.

Adapted vs. frozen encoder32.40 vs. 23.41

PSNR

MLP vs. UNet decoder32.40 vs. 29.58

PSNR

Additive conditioning32.40

31.93 AdaLN · 22.66 cross-attn

Method · 02

Retrieval-augmented retouching

With multiple reference pairs, a user's preferred adjustment may vary with the image. Averaging every reference into one style ignores this content-dependent behavior.

RAR retrieves the top-K reference inputs most similar to the query and forms a similarity-weighted combination of their retouching latents. The adapted SigLIPv2 features encode retouching-relevant color and tone, which makes them suitable for this retrieval.

Retrieval comparison using adapted and frozen SigLIPv2 features
The adapted encoder retrieves images with similar color and tone, while frozen SigLIPv2 retrieval is mainly semantic.
01

Multiple references

On PPR10K-Groups, performance improves consistently as the number of references increases.

1 reference22.51
3 references23.92
6 references25.27
Multi-reference PPR10K-Groups comparison
02

Standard benchmarks

AdaInt and RSFNet are trained on each user's 100 references. RefRetouch uses the references at inference without user-specific training.

Dataset · 100 refsMSMAdaInt*RSFNet*RefRetouch
MIT-FiveK22.05 / .80722.49 / .85422.66 / .86023.34 / .920
PPR10K20.06 / .87322.82 / .93222.22 / .92322.47 / .942

PSNR ↑ / SSIM ↑ · *user-specific training

03

RAR ablation

With 100 MIT-FiveK references, retrieving the five most relevant examples is better than combining all references.

Top-121.94
Top-323.34
Top-523.82
All21.94

Efficiency and application

Fast retouching and color-tone retrieval

MLP to 3D-LUT

The global pixel-wise MLP is a deterministic RGB-to-RGB function for a fixed latent. Evaluating it on a regular RGB grid converts it into a 3D-LUT.

1.72 mslatency at 1024²
622.7images/s at 1024²
1.98 GBpeak memory at 1024²
25.28 dBPPR10K-Groups · LUT

MLP: 43.21 ms, 23.0 images/s, 6.93 GB, and 25.27 dB. The LUT provides the speedup with negligible quality difference.

Color-and-tone retrieval

The adapted encoder can retrieve images by color and tone, which can support dataset preparation and image organization.

Color-and-tone image retrieval using the adapted RefRetouch encoder

Citation

RefRetouch

BibTeX
@inproceedings{weldengus2026refretouch,
  title     = {RefRetouch: Personalized Image Retouching
               without Test-time Fine-tuning},
  author    = {Weldengus, Temesgen Muruts and Liu, Binnan and
               Kou, Fei and Lyu, Youwei and Chen, Jinwei and
               Fan, Qingnan and Zou, Changqing},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}

References

Related work and benchmarks

  1. C. Yang et al. “AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image Enhancement.” CVPR, 2022.
  2. W. Ouyang et al. “RSFNet: A White-box Image Retouching Approach Using Region-specific Color Filters.” ICCV, 2023.
  3. Z.-P. Duan et al. “DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts.” AAAI, 2025.
  4. V. Bychkovsky et al. “Learning Photographic Global Tonal Adjustment with a Database of Input / Output Image Pairs.” CVPR, 2011.
  5. J. Liang et al. “PPR10K: A Large-scale Portrait Photo Retouching Dataset with Human-region Mask and Group-level Consistency.” CVPR, 2021.
  6. S. Kosugi and T. Yamasaki. “Personalized Image Enhancement Featuring Masked Style Modeling.” IEEE TCSVT, 2024.
  7. Y. Song, H. Qian, and X. Du. “StarEnhancer: Learning Real-time and Style-aware Image Enhancement.” ICCV, 2021.
  8. M. Tschannen et al. “SigLIP 2: Multilingual Vision-language Encoders with Improved Semantic Understanding, Localization, and Dense Features.” arXiv, 2025.