Diffusion samplers implemented from scratch on top of pretrained Stable Diffusion weights, used to measure where text-guided image editing stops working.
diffusers is used only as a weight loader. The denoising loop, the
noise schedule, the timestep grid, and classifier-free guidance are all
implemented here and checked against the reference implementation.
| Pipeline | Initial latent | Conditioning | Guidance |
|---|---|---|---|
| text2img | pure noise | text | 2-way CFG |
| SDEdit | image latent + noise at t0 | text | 2-way CFG |
| InstructPix2Pix | pure noise | text + image latent (channel concat) | 3-way |
All three pipelines share a single denoising loop (src/difflab/samplers/base.py::denoise).
SDEdit has no usable operating point: for a colour edit, the prompt is
first recognised at strength 0.5 -- the same strength at which the
source image already starts to collapse. InstructPix2Pix holds up for
global and attribute edits at image_scale >= 1.5, but fails object
removal in all 12 grid cells tested -- it either deletes nothing or
deletes the entire room.
Full write-up with images and a two-layer root-cause analysis (training distribution vs. interface): docs/baseline-ko.md (Korean). The follow-up experiment that disambiguates the two hypotheses -- masking the edit region to test whether the interface, not the training distribution, was the blocker: docs/masking-ko.md (Korean).
webui/ runs a small local server (models stay resident in the
process, loaded lazily per pipeline) for generating anchor images and
trying SDEdit / InstructPix2Pix against them interactively, including
live parameter sweeps rendered as a slider. See
webui/README.md.
See SETUP.md.
See NOTES.md.