Skip to content

Using the perceptual / semantic compression boundary to design latent spaces #101

Description

@dltmdwoOS

In the ‘Perceptual and Semantic Compression in Latent Diffusion Models’ slide, the Rate–Distortion curve suggests that below a certain rate the distortion increases sharply, which I interpret as the region where semantic information starts to be lost. In the context of Latent Diffusion Models, when tuning the auto-encoder’s compression strength (downsampling factor, number of latent channels, strength of KL/VQ regularization ...), is there any way to quantitatively estimate or exploit this boundary between the ‘semantic’ and ‘perceptual’ regions on the curve? For example, are there procedural guidelines or experimental results that measure the R–D curve on a dataset and then choose the latent dimensionality near the critical point where semantic distortion begins to rise?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions