Understand the preference
Encode the retouching difference in the reference pairs as a retouching latent.
ECCV 2026

Motivation
Users have different retouching preferences. Prior methods such as AdaInt [1], RSFNet [2], and DiffRetouch [3] learn a fixed style from one or several experts' paired data. A new user or style requires additional data and retraining.
Personalized retouching should instead adapt to a user on the go from a few exemplars that show the user's preference: tuning-free image retouching personalization.
Encode the retouching difference in the reference pairs as a retouching latent.
Select and combine the relevant preference exemplars for each new image.
Method · 01
A LoRA-adapted Siamese SigLIPv2 encoder [8] processes the original and retouched images and represents their color and tone transformation as a retouching latent.
A lightweight conditional MLP maps each original RGB value to its retouched RGB value. Because this decoder operates only in color space, the latent must carry the retouching transformation rather than image content.

Results · Retouching latent
RefRetouch reconstructs the paired target more faithfully than MSM [6] on all three datasets.
| Dataset | MSM | MSM retrained | RefRetouch |
|---|---|---|---|
| VCIRB | 27.28 / .900 / .099 | 29.61 / .923 / .069 | 32.42 / .974 / .035 |
| MIT-FiveK | 28.94 / .851 / .137 | 26.32 / .803 / .199 | 32.95 / .972 / .028 |
| PPR10K-A | 29.28 / .943 / .075 | 29.71 / .936 / .086 | 34.14 / .984 / .019 |
PSNR ↑ / SSIM ↑ / LPIPS ↓
One reference latent transfers the retouching style to unseen images. On VCIRB, RefRetouch reaches 29.13 dB PSNR, compared with 24.35 dB for retrained MSM and 21.00 dB for StarEnhancer [7].

Task adaptation, a color-space MLP, and additive conditioning give the strongest reconstruction.
PSNR
PSNR
31.93 AdaLN · 22.66 cross-attn
Method · 02
With multiple reference pairs, a user's preferred adjustment may vary with the image. Averaging every reference into one style ignores this content-dependent behavior.
RAR retrieves the top-K reference inputs most similar to the query and forms a similarity-weighted combination of their retouching latents. The adapted SigLIPv2 features encode retouching-relevant color and tone, which makes them suitable for this retrieval.

On PPR10K-Groups, performance improves consistently as the number of references increases.

AdaInt and RSFNet are trained on each user's 100 references. RefRetouch uses the references at inference without user-specific training.
| Dataset · 100 refs | MSM | AdaInt* | RSFNet* | RefRetouch |
|---|---|---|---|---|
| MIT-FiveK | 22.05 / .807 | 22.49 / .854 | 22.66 / .860 | 23.34 / .920 |
| PPR10K | 20.06 / .873 | 22.82 / .932 | 22.22 / .923 | 22.47 / .942 |
PSNR ↑ / SSIM ↑ · *user-specific training
With 100 MIT-FiveK references, retrieving the five most relevant examples is better than combining all references.
Efficiency and application
The global pixel-wise MLP is a deterministic RGB-to-RGB function for a fixed latent. Evaluating it on a regular RGB grid converts it into a 3D-LUT.
MLP: 43.21 ms, 23.0 images/s, 6.93 GB, and 25.27 dB. The LUT provides the speedup with negligible quality difference.
The adapted encoder can retrieve images by color and tone, which can support dataset preparation and image organization.

Citation
@inproceedings{weldengus2026refretouch,
title = {RefRetouch: Personalized Image Retouching
without Test-time Fine-tuning},
author = {Weldengus, Temesgen Muruts and Liu, Binnan and
Kou, Fei and Lyu, Youwei and Chen, Jinwei and
Fan, Qingnan and Zou, Changqing},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}References