Note
This is a weekend-ish project I worked on for fun.
I want to thank the authors of Guiding a Diffusion Model with a Bad Version of Itself for providing the valuable test bench I evaluated my model against
What if we trained a diffusion model capable of estimating the confidence of each prediction?
I also decided to follow these tight constraints:
- Maximize the log probability and nothing else.
- The prediction distribution must be Gaussian.
- The result should generalize the standard diffusion model framework.
- If the confidence level is set equal to the diffusion process noise
$\sigma$ , it should reduce exactly to the equations from EDM.
So, I did the math and ran a small-scale experiment based on Guiding a Diffusion Model with a Bad Version of Itself to compare my method to the “standard” diffusion modeling approach.
In this small-scale experiment, the two models ended up with nearly identical performance. More experiments on complex datasets will be needed to determine whether the method brings real improvements.
Left: standard EDM method. Right: my confidence-aware method.
The code is lightweight and can even run on a laptop.
To run both experiments (plus the score-matching experiment), simply run the ablations.py script.
-
Install
uvwith pip:pip install uv
-
Create and activate a virtual environment:
uv venv source .venv/bin/activate # or .venv\Scripts\activate on Windows
-
Install the package:
uv pip install -e . -
Run the experiment:
python ablations.py
Results will be saved to a folder.
Warning
This is the nerd area, proceed with caution
Let’s make score-based diffusion modeling more rigorous.
Suppose we have a dataset containing a single sample
This implies that the distribution is:
But what if
Which gives:
You might think these two processes need different treatments — but we can hit two birds with one stone.
From now on, we consider the general case
And we allow
How do you choose the loss function? You don’t — nature does it for you.
Given a data point
We then minimize the expected negative log-likelihood:
Which decomposes into:
Carrying out the KL divergence, we get:
At the minimum, the optimal variance satisfies:
So the learned variance captures both the added noise and the model’s prediction error — an adaptive uncertainty estimate.
Also, the gradient with respect to
This resembles the gradient of MSE, but weighted by the predicted error — providing a natural form of loss weighting. This is interesting because in EDM2, a similar weighting was manually engineered, but here it emerges on its own from the equations 🔥.
In score-based diffusion, we aim to estimate the score function:
Using Tweedie's formula, a natural estimator of the score is:
So we don’t need
We can solve the denoising ODE using the usual techniques:
When implementing this in practice, we must be cautious — neural networks don't handle very large or very small values well.
Training a model to predict
Where
We now substitute everything back into the original loss:
After substitution, we get a numerically stable loss for both
[!Hint] To verify this, first substitute
$\sigma_\phi$ and then$\mu_\theta$ into the original loss.
Using this reparametrization, we can rewrite the log-probability as:
This formulation is numerically stable for all values of