FreeEyeglass: Training-free and Target-mask-free
Eyeglass Transfer for Facial Videos

TMLR 2026
1The University of Osaka, 2CyberAgent AI Lab

FreeEyeglass transfers any eyeglasses to facial videos in a training-free way.

Abstract

The rise of e-commerce and short-video platforms has fueled demand for realistic video-based virtual try-on. Unlike virtual try-on of clothing, which has been actively studied to date, virtual try-on of eyeglasses is uniquely challenging: they align closely with facial structure and strongly affect facial identity, making the faithful preservation of unedited regions especially important. Existing generative editing approaches, such as GAN- and diffusion-based methods, lack reconstruction objectives and often rely on inpainting, which fails to ensure identity consistency. We argue that semantic editing requires not only plausible generation but also faithful reconstruction, making autoencoder-based latent spaces a natural fit. We introduce a training-free, reference-guided framework for video eyeglass transfer built on Diffusion Autoencoders (DiffAE). By blending semantic features in the encoder and incorporating spatial-temporal self-attention, our method achieves realistic, identity-preserving, and temporally consistent results, and points to the potential of autoencoder-based latent spaces for local video editing.

Results

Example 1 reference eyeglasses

Reference

Target

Ours

Example 2 reference eyeglasses

Reference

Target

Ours

Example 3 reference sunglasses

Reference

Target

Ours

Method Overview

Overview of the FreeEyeglass pipeline
Overview of our FreeEyeglass pipeline. Given a target video (i.e., target image sequence) and a reference image of desired eyeglasses, we blend the features of the reference eyeglasses and the target image sequence to obtain a blended semantic latent sequence. We then compute the stochastic latent sequences with the input images and semantic latent sequences through conditional DDIM inversion and construct a blended stochastic latent sequence. Using the blended stochastic latent and semantic latent sequences, we can obtain the final edited image sequence with our desired eyeglasses semantically and naturally placed through condition DDIM sampling.

More Results

Citation

@article{chan2026freeeyeglass,
  title={FreeEyeglass: Training-free and Target-mask-free Eyeglass Transfer for Facial Videos},
  author={Chan, Weng Ian and Huang, Yuantian and Yang, Xingchao and Okura, Fumio and Taketomi, Takafumi},
  journal={Transactions on Machine Learning Research},
  year={2026},
  url={https://openreview.net/forum?id=6aFRoQcm3H}
}