Latent diffusion models rely on powerful pretrained Variational Autoencoders (VAEs) to project images into a compact latent space where diffusion sampling is both expressive and computationally efficient. We demonstrate that this same latent representation can be exploited to build a practical video compressor whose rate–distortion performance surpasses the industrial H.265/HEVC standard across a wide range of bitrates. The proposed pipeline first encodes a sparse set of keyframes with the latent-diffusion VAE, after which the resulting latent vectors are entropy-coded to remove residual spatial redundancy and drive the bitrate down. Starting from these decoded keyframes, each Group-of-Pictures is reconstructed entirely in latent space through a diffusion-based temporal interpolation process that synthesizes intermediate frames conditioned on their temporal context. To push performance further, we introduce two complementary fine-tuning strategies that freeze the decoder while learning frame-specific latent codes: a zero-shot configuration that operates with the off-the-shelf model and an adaptive configuration that refines both the latent codes and a lightweight subset of encoder parameters for every input video. Extensive experiments on standard benchmarks confirm that, under all tested conditions, our method consistently delivers higher objective quality and better perceptual fidelity than H.265 at comparable or lower bitrates.
Diffusion Autoencoders are Foundation Video Compressors / Niccoli, N., Galteri, L., Seidenari, L.. - ELETTRONICO. - (2025), pp. 3964-3972. (2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) ).
Diffusion Autoencoders are Foundation Video Compressors
Niccoli, Niccolo'
;Galteri, Leonardo;Seidenari, Lorenzo
2025
Abstract
Latent diffusion models rely on powerful pretrained Variational Autoencoders (VAEs) to project images into a compact latent space where diffusion sampling is both expressive and computationally efficient. We demonstrate that this same latent representation can be exploited to build a practical video compressor whose rate–distortion performance surpasses the industrial H.265/HEVC standard across a wide range of bitrates. The proposed pipeline first encodes a sparse set of keyframes with the latent-diffusion VAE, after which the resulting latent vectors are entropy-coded to remove residual spatial redundancy and drive the bitrate down. Starting from these decoded keyframes, each Group-of-Pictures is reconstructed entirely in latent space through a diffusion-based temporal interpolation process that synthesizes intermediate frames conditioned on their temporal context. To push performance further, we introduce two complementary fine-tuning strategies that freeze the decoder while learning frame-specific latent codes: a zero-shot configuration that operates with the off-the-shelf model and an adaptive configuration that refines both the latent codes and a lightweight subset of encoder parameters for every input video. Extensive experiments on standard benchmarks confirm that, under all tested conditions, our method consistently delivers higher objective quality and better perceptual fidelity than H.265 at comparable or lower bitrates.| File | Dimensione | Formato | |
|---|---|---|---|
|
Niccoli_Diffusion_Autoencoders_are_Foundation_Video_Compressors_ICCVW_2025_paper.pdf
accesso aperto
Tipologia:
Pdf editoriale (Version of record)
Licenza:
Open Access
Dimensione
733.53 kB
Formato
Adobe PDF
|
733.53 kB | Adobe PDF |
I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



