Hybrid Automatic Video Colorizer (HAVC) server that exposes a GPU-accelerated colorization pipeline for black-and-white images and video frames based on Diffusion Transformer (DiT) models.
4 backends, one API : pick the one that fits your hardware:
- nunchaku-qwen: SVDQuant FP4/INT4 transformer via Nunchaku : 4 sec/frame¹, requires RTX 30/40/50 (16GB+ VRAM, 64GB RAM) & CUDA 13.0
- gguf-qwen: ComfyUI-native GGUF pipeline (Q3_K_S, Q4_K_S, Q5_K_M, Q6_K, Q8_0) : 12 sec/frame², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), zero ComfyUI GUI dependency
- longcat-gguf: LongCat-Image-Edit-Turbo GGUF pipeline (Q3_K_M–Q8_0) : ~12 sec/frame², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), better image quality than gguf-qwen, zero ComfyUI GUI dependency
- qwen21-viggle: Qwen-Image-2.1 (native ComfyUI int8 ConvRot UNet) + Viggle-Turbo LoRA : ~6 sec/frame¹ (Fast Pipeline) or ~8-11 sec/frame, runs on RTX 30/40/50 (14GB+ VRAM, 32GB+ RAM), optional
enhance_prompt(Qwen3-VL image-aware prompt rewriting)
¹ Measured with Fast Pipeline (paired inference). ²
gguf-qwen/longcat-ggufdon't support paired inference — single-image time, not directly comparable. Speed/hardware/quality trade-offs, the recommended picks and the full measurement notes: docs/backends.md.
Installer (recommended) — download HAVC-Setup-<version>.exe from the Releases page and run it: a single self-contained file (no .NET runtime needed) that installs the pinned Python runtime, the server stack, the GUI, external tools and the vs-cmnet2 plugins/weights — and keeps everything updated (Check for updates, Repair, Uninstall from the manager). Model weights are downloaded on first use and preserved across updates. Details: docs/installation.md.
⚠️ Windows SmartScreen will warn on first run because the exe is not code-signed: More info → Run anyway. The sha256 of every asset is published in the release notes.
Manual installation (advanced) — clone the repository and run install.cmd (requires Git + Python 3.12): step-by-step instructions in docs/installation.md.
From the install folder (or the repository root on a manual install):
HAVC.vbs # desktop GUI (the recommended front-end)
HAVC-Server.cmd # console server — pick the model at the prompt
start_server.cmd q3 # console server with a backend/quantization or a config name
CLI arguments, the full launch-script reference and the suggested inference steps: docs/server_usage.md.
- Windows 10/11, NVIDIA RTX 30/40/50 GPU, CUDA 13.0+
- VRAM / RAM by backend: nunchaku-qwen 16 GB+ / 64 GB+ · gguf-qwen and longcat-gguf 12 GB+ / 32 GB+ · qwen21-viggle 14 GB+ / 32 GB+ — details
- Disk: a few GB for the software stack; model weights (up to tens of GB for the largest configs) are downloaded on first use
- 📦 4 backends, one API : nunchaku-qwen (FP4/INT4, 4 sec/frame) for speed, gguf-qwen and longcat-gguf (Q3, …, Q8, 12 sec/frame) for lower VRAM, qwen21-viggle (int8 ConvRot UNet, ~8-11 sec/frame) with optional Qwen3-VL prompt rewriting
- 🎨 Batch colorization : process entire directories of B&W images via filesystem paths
- 🖼️ Paired inference : colorize two images in a single forward pass (faster, temporally consistent)
- 📡 In-memory RPC : pass raw PNG frames over XML-RPC without touching the filesystem (ideal for video pipelines)
- ⚡ 4-step lightning model : SVDQuant FP4 quantized transformer for maximum throughput
- 🔒 Thread-safe : pipeline loading and stop control are protected by locks; every RPC call runs in its own thread
- ⚙️ Startup preload : optional
--load-pipelineflag loads the model at boot from a JSON config file - 🚀 Shared memory transport : zero-copy image transfer for same-host deployments (~23% faster than standard RPC)
- Installation — installer, manual setup, updating, project layout
- Backends & requirements — choosing the right backend
- Server usage — server startup, launch scripts, CLI arguments, suggested steps
- Pipeline configuration — config files and key reference
- RPC API — XML-RPC API, example clients, shared-memory transport
- GUI usage — the desktop client
- Troubleshooting
- What's New — changelog
- Model: Qwen/Qwen-Image-Edit-2511, Qwen/Qwen-Image-2.1, LongCat-Image-Edit-Turbo
- VapourSynth filter for video colorization with CMNET2: vs-cmnet2
- Viggle-Turbo LoRA: Viggle/Qwen-Image-2.1-viggle-turbo
- Nunchaku quantization: Nunchaku / SVDQuant
- GGUF dequantization kernels: adapted from ComfyUI-GGUF (Apache 2.0), Qwen3-VL mmproj support from the ComfyUI-GGUF-Reboot fork (molbal)
- Pipeline: Hugging Face Diffusers