Skip to content

Repository files navigation

difflab

Diffusion samplers implemented from scratch on top of pretrained Stable Diffusion weights, used to measure where text-guided image editing stops working.

diffusers is used only as a weight loader. The denoising loop, the noise schedule, the timestep grid, and classifier-free guidance are all implemented here and checked against the reference implementation.

What's implemented

Pipeline Initial latent Conditioning Guidance
text2img pure noise text 2-way CFG
SDEdit image latent + noise at t0 text 2-way CFG
InstructPix2Pix pure noise text + image latent (channel concat) 3-way

All three pipelines share a single denoising loop (src/difflab/samplers/base.py::denoise).

Finding

SDEdit has no usable operating point: for a colour edit, the prompt is first recognised at strength 0.5 -- the same strength at which the source image already starts to collapse. InstructPix2Pix holds up for global and attribute edits at image_scale >= 1.5, but fails object removal in all 12 grid cells tested -- it either deletes nothing or deletes the entire room.

Full write-up with images and a two-layer root-cause analysis (training distribution vs. interface): docs/baseline-ko.md (Korean). The follow-up experiment that disambiguates the two hypotheses -- masking the edit region to test whether the interface, not the training distribution, was the blocker: docs/masking-ko.md (Korean).

Interactive viewer

webui/ runs a small local server (models stay resident in the process, loaded lazily per pipeline) for generating anchor images and trying SDEdit / InstructPix2Pix against them interactively, including live parameter sweeps rendered as a slider. See webui/README.md.

Setup

See SETUP.md.

Notes

See NOTES.md.

About

Diffusion samplers implemented from scratch, used to measure where text-guided image editing stops working

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages