Skip to content

Latest commit

Β 

History

249 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

ComfyUI-FFMPEGA

The ultimate video editing suite for ComfyUI β€” edit with natural language or hands-on manual controls.

ComfyUI Version License Dependencies Last Commit Activity

FFMPEGA Showcase

Use AI to describe edits in plain English, or take full manual control with the Effects Builder and text presets β€” no LLM required.

Features β€’ Examples β€’ Installation β€’ Quick Start β€’ Prompt Guide β€’ Skills β€’ LLM Setup β€’ Troubleshooting β€’ Contributing β€’ Changelog


πŸš€ What's New in v2.21.0

See CHANGELOG.md for the full detail.

  • πŸ”οΈ Marigold V2 β€” depth, surface normals and albedo from a frozen Qwen-Image-Edit-2509 DiT plus a per-task LoRA, VAE and baked conditioning. One deterministic Euler step, no CFG, no seed. Reached through depth (v2) / normals (v2) / albedo (v2) on the existing marigold mode, so nothing already in your workflows changes; also selectable as a depth_backend on the Shader Overlay node. ~22.5 GB for the first task, ~2 GB per task after (the base is shared), at roughly 3–4 s per frame
  • 🎚️ Fixes ComfyUI's stock Marigold V2 sampling β€” the shipped blueprints patch ModelSamplingAuraFlow with sampling="flow", which hands the DiT a half-scale latent and returns 0.5Β·(z βˆ’ v). This backend defaults to IMG_TO_IMG_VELOCITY, which yields z βˆ’ v from an unscaled latent and matches the reference implementation exactly; flow stays selectable to reproduce the stock template
  • πŸ”§ BlockSwap block discovery β€” the swap budget was sized from model.diffusion_model.blocks only, so any model naming its stack otherwise fell back to an estimate scaled by Wan's 40-block depth. It now also finds transformer_blocks, double_blocks and single_blocks, which matters for 60-block Qwen-Image
πŸ“‹ Previous Releases
Version Highlights
v2.20.0 SCAIL-2 pose-driven character animation, FlashVSR one-step super-resolution, SVI 2.0 Pro infinite-length video, Wan-Animate, MatAnyone2 matting, Meta Sapiens2, SHARP, PhyFPS, multi-source comparison, corrected BT.709 video levels
v2.19.0 DreamID-Omni talking-head generation, FaceCam camera control, Fish Speech TTS, Foundation-1 music samples, Frame Picker node, 15 new GLSL shaders (70 total)
v2.18.0 Kiwi-Edit AI video editing, SAM3 + Kiwi-Edit, RTX Video Super Resolution, SeedVR AI Upscaling, FacePoke expression presets
v2.17.0 FacePoke interactive face editor, driving video reference, shader effects system, Flux Klein FP8, onion skin compositing, unified audio output mode
v2.16.0 ACE-Step AI music generation, SAM-Audio source separation, Video Editor v2 (10 panels), AudioX vocal enhancement, NormalCrafter, Video Depth Anything
v2.15.0 MiniMax-Remover, 5 new no-LLM modes, auto-VRAM tile sizing, VRAM management overhaul
v2.14.0 Video Editor NLE node with timeline, razor, crop, transitions, text overlays, keyboard shortcuts
v2.13.0 AI Background Removal (BRIA RMBG), FLUX Klein toggle, Edit FFmpeg fallback, smarter defaults
v2.12.0 AI Face Animation (LivePortrait), MMAudio in-process inference, MCP progressive disclosure, LaMa safetensors conversion
v2.11.0 MMAudio in-process migration, generate_audio no-LLM mode, MCP tools, CLI binary caching
v2.10.0 FLUX Klein in-process migration, interactive mask drawing UI, model output caching
v2.9.1 AI object removal & editing (FLUX Klein 4B), AI audio generation (MMAudio), AI lip sync (MuseTalk), modular architecture refactor
v2.8.0 Effects Builder node, manual no-LLM mode, text node presets, SAM3 subprocess isolation, 15+ bug fixes
v2.7.0 SAM3 auto-mask & greenscreen, LaMa inpainting, programmatic tool calling (PTC)
v2.6.5 Whisper auto-transcription, karaoke subtitles, whisper model/device controls
v2.6.0 HandlerResult contract, compose decomposition, TextInput node, PiP audio mixing, CLI retry
v2.5.0 PiP borders, Ollama VL auto-embedding, overlay animation delegation
v2.4.0 Zero-memory image paths, pipeline chaining fixes, handler module extraction
v2.3.0 Token usage tracking, LUT color grading, vision system, audio analysis
v2.2.0 200 skills, dynamic input slots, 48 new skills across all categories
v2.0.0 Dynamic input slots, concat, xfade, split screen, animated overlays, text overlays
v1.9.0 Claude CLI, Cursor Agent CLI, Qwen CLI connectors, 50+ context menu presets

πŸ“„ See CHANGELOG.md for the complete version history.


NotebookLM Overview: Exploring the features and capabilities of ComfyUI-FFMPEGA. (Click to watch on YouTube)

✨ Features

πŸ—£οΈ Natural Language Editing

Describe edits in plain text: "Make it cinematic with a fade in", "Speed up 2x", "VHS look with grain". The AI agent interprets your prompt and builds the FFMPEG pipeline automatically.

πŸ—οΈ Manual Mode β€” No AI Required

Use the Effects Builder to visually compose up to 5 effects with parameters. Add text overlays and subtitles via preset-powered Text nodes. Full editing control with zero LLM dependency.

πŸ€– Multi-LLM Support

Works with Ollama (local, free) and CLI tools (Gemini CLI, Claude Code, Cursor Agent, Qwen Code). Use any local model β€” Llama 3.1, Qwen3, Mistral, and more. Or skip the LLM entirely.

🎨 200+ Skills

200+ video editing skills across visual effects, audio processing, spatial transforms, temporal edits, encoding, cinematic presets, vintage looks, social media, creative effects, text animations, editing & composition, audio visualization, multi-input operations, transitions, concat, split screen, and AI-powered skills (DreamID-Omni talking-head generation, FaceCam camera control, Fish Speech TTS, Foundation-1 music samples, ACE-Step music generation, SAM-Audio source separation, AudioX vocal enhancement, Whisper transcription, SAM3 masking, MiniMax-Remover object removal, MMAudio generation, MuseTalk lip sync, LivePortrait face animation, NormalCrafter surface normals, Video Depth estimation, AI Upscaling, Marigold dense vision).

🎨 Right-Click Presets

26 built-in Effects Builder presets and 10 Text node presets with example content. Save/load/delete your own custom presets. One-click clear to reset.

⚑ Batch & Preview

Process multiple videos with the same instruction. Generate quick low-res previews before committing to full renders. Quality presets from draft to lossless.


🎬 Examples

Worked examples: prompts, the pipelines they produce, and the resulting output.

β†’ docs/EXAMPLES.md

πŸ“¦ Installation

Requirements

  • ComfyUI (latest)
  • Python 3.10+
  • FFMPEG installed and in PATH (install guide)
  • Node.js 18+ (required for CLI tools: Gemini CLI, Claude CLI, Qwen CLI β€” download)
  • Ollama (optional, for local LLM inference β€” download)

Option 1: ComfyUI Manager (Recommended)

  1. Open ComfyUI Manager
  2. Search for ComfyUI-FFMPEGA
  3. Click Install

Option 2: Manual Install

Linux / macOS
cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/AEmotionStudio/ComfyUI-FFMPEGA.git
cd ComfyUI-FFMPEGA
pip install -r requirements.txt
Windows (PowerShell)
cd C:\path\to\ComfyUI\custom_nodes
git clone https://github.com/AEmotionStudio/ComfyUI-FFMPEGA.git
cd ComfyUI-FFMPEGA
pip install -r requirements.txt

Note: Use whichever Python package manager your ComfyUI venv uses (pip, uv pip, etc.). The above commands assume pip is available in your ComfyUI virtual environment.

Restart ComfyUI after installation.


πŸš€ Quick Start

  1. Add an FFMPEG Agent node to your workflow
  2. Connect a video path or use the input field
  3. Enter a natural language prompt
  4. Select your LLM model
  5. Run the workflow

πŸ’‘ Tip: Right-click the FFMPEG Agent node to open the FFMPEGA Presets context menu β€” 200+ categorized effects you can apply with a single click, no prompt typing needed. Great for quick edits or discovering what's available.

Example Prompts

Prompt What It Does
"Make it cinematic with a vignette" Adds letterbox, color grade, and edge darkening
"Speed up 2x, keep the audio pitch" Doubles speed with pitch-corrected audio
"Make it look like old VHS footage" Adds noise, color shift, scan lines
"Trim first 5 seconds, resize to 720p" Cuts intro and scales down
"Underwater look with echo on audio" Blue tint, blur, and audio echo
"Pixelate it like an 8-bit game" Mosaic/pixel art effect
"Add 'Subscribe!' text at the bottom" Text overlay with positioning
"Cyberpunk style with neon glow" High-contrast neon aesthetic
"Normalize audio, compress for web" Loudness normalization + web optimization
"Spin the video clockwise" Continuous animated rotation
"Add camera shake" Random shake/earthquake effect
"Fade in from black and out at the end" Smooth intro/outro transitions
"Wipe reveal from the left" Directional wipe reveal animation
"Make it pulse like a heartbeat" Rhythmic zoom breathing effect
"Show audio waveform at the bottom" Audio visualization overlay
"Arrange these images in a grid" Multi-image grid collage
"Create a slideshow with fades" Image slideshow with transitions
"Overlay the logo in the corner" Picture-in-picture / watermark
"Create a side-by-side comparison" Video next to image in 2-column grid
"Create a slideshow starting with the video" Video first, then image slides
"Overlay images in the corners" Multiple images auto-placed in corners
"Split screen, use audio from audio_b" Side-by-side video with specific audio track
"Split screen, mix both audio tracks" Side-by-side with both audio tracks blended

πŸ’¬ Prompt Guide

When using an LLM, the AI agent interprets your natural language and maps it to skills with specific parameters. Here's how to get the best results. (For manual editing without an LLM, see the Effects Builder and Text Input sections.)

Specifying Exact Values

You can request specific parameter values and the agent will use them directly:

Prompt What the Agent Does
"Set brightness to 0.3" brightness:value=0.3
"Blur with strength 20" blur:radius=20
"Speed up to 3x" speed:factor=3.0
"Crop to 1280x720" crop:width=1280,height=720
"Deband with threshold 0.3 and range 32" deband:threshold=0.3,range=32
"CRF 18, slow preset" quality:crf=18,preset=slow
"Fade in for 3 seconds" fade:type=in,duration=3

What Works Well βœ…

  • Explicit numbers: "brightness 0.2", "speed 1.5x", "CRF 20" β€” the agent maps these directly
  • Named presets: "VHS look", "cinematic style", "noir" β€” triggers multi-step preset pipelines
  • Chaining operations: "Trim first 5 seconds, resize to 720p, add vignette" β€” executes in order
  • Descriptive goals: "Make it look warmer", "Remove the green screen" β€” the agent picks the right skills
  • Technical terms: "denoise", "deband", "normalize audio" β€” maps to exact FFmpeg filters

What Might Not Work as Expected ⚠️

  • Vague intensity words: "Make it very blurry" or "a little brighter" β€” the agent has to guess what number "very" or "a little" means. Tip: use a specific value instead: "blur with radius 15"
  • Out-of-range values: Parameters are auto-clamped to their valid range. If you ask for "brightness 5.0" it caps at the max (1.0)
  • Complex compositing: Multi-layer effects with precise timing may need to be broken into separate passes
  • Format-dependent features: Some effects (like transparency) require specific output formats. H.264/MP4 doesn't support alpha channels

🧠 AI Models (Auto-Downloaded)

Some skills use AI models that auto-download on first use. You can disable automatic downloads with the allow_model_downloads toggle on the FFMPEG Agent node β€” runs requiring a missing model will fail with a clear message and a manual download link.

All models are mirrored to first-party AEmotionStudio HuggingFace repos for supply chain resilience. Downloads try the AEmotionStudio mirror first, then fall back to upstream sources.

Model Size Stored In Triggered By Manual Download
SAM3 (Segment Anything 3) ~300 MB ComfyUI/models/SAM3/ auto_mask skill, sam3_masking no-LLM mode, Effects Builder SAM3 target AEmotionStudio/sam3 β€” download sam3.safetensors
SAM3.1 (Multiplex Tracker) ~3.5 GB ComfyUI/models/SAM3.1/ Same triggers with sam_version = sam3.1 (default). Video masking runs in-process on ComfyUI's native SAM3 model (VRAM-managed, ~3 GB peak); set FFMPEGA_SAM3_NATIVE=0 to force the legacy subprocess path AEmotionStudio/sam3.1 β€” download sam3.1_multiplex.safetensors
Whisper large-v3 ~3 GB ComfyUI/models/whisper/ auto_transcribe, karaoke_subtitles skills, transcribe / karaoke_subtitles no-LLM modes AEmotionStudio/whisper-models
Whisper medium ~1.5 GB ComfyUI/models/whisper/ Same as above (set whisper_model to medium) Same as above
Whisper small ~500 MB ComfyUI/models/whisper/ Same as above (set whisper_model to small) Same as above
Whisper base ~150 MB ComfyUI/models/whisper/ Same as above (set whisper_model to base) Same as above
Whisper tiny ~75 MB ComfyUI/models/whisper/ Same as above (set whisper_model to tiny) Same as above
LaMa (Large Mask Inpainting) ~195 MB ~/.cache/torch/hub/checkpoints/ auto_mask:effect=remove (legacy fallback) AEmotionStudio/lama-inpainting β€” download big-lama.safetensors
FLUX Klein 4B (Editing/Removal) ~15 GB (bf16) ComfyUI/models/flux_klein/ auto_mask:effect=remove, auto_mask:effect=edit AEmotionStudio/flux-klein
MiniMax-Remover (Object Removal) ~2.5 GB ComfyUI/models/minimax_remover/ auto_mask:effect=remove (when use_minimax_remover is On) AEmotionStudio/minimax-remover
MMAudio (Video-to-Audio) ~5.5 GB ComfyUI/models/mmaudio/ generate_audio skill AEmotionStudio/mmaudio-models
MuseTalk (Lip Sync) ~1.6 GB (fp16) ComfyUI/models/musetalk/ lip_sync skill AEmotionStudio/musetalk-models
LivePortrait (Face Animation) ~497 MB ComfyUI/models/liveportrait/ animate_portrait skill, animate_portrait no-LLM mode AEmotionStudio/liveportrait-models
Video Depth Anything (Temporal Depth) ~102–670 MB ComfyUI/models/video_depth/ video_depth no-LLM mode AEmotionStudio/video-depth-anything
Marigold v1.1 (Dense Vision) ~2.5 GB per mode Auto-downloaded by diffusers marigold no-LLM mode (depth/normals/appearance/lighting) AEmotionStudio/marigold-depth-v1-1
Marigold V2 (Dense Vision) ~22.5 GB first task, ~2 GB each after (20.5 GB base is shared) ComfyUI/models/{diffusion_models,loras,vae,embeddings}/ marigold no-LLM mode, depth (v2) / normals (v2) / albedo (v2); also a depth_backend option on the Shader Overlay node Comfy-Org/marigold-v2-0
AI Upscaler (Real-ESRGAN / HAT / DAT / SwinIR) ~17–170 MB per model ComfyUI/models/upscale_models/ ai_upscale skill, ai_upscale no-LLM mode AEmotionStudio/ai-upscale-models
BRIA RMBG (rembg) ~270 MB ~/.u2net/ remove_background skill Install with pip install 'comfyui-ffmpega[masking]' β€” model auto-fetched by rembg
ACE-Step 1.5 (Music Generation) ~5 GB ComfyUI/models/acestep/ ace_step no-LLM mode, generate_music skill AEmotionStudio/ACE-Step
SAM-Audio (Source Separation) ~1.2 GB (large) / ~600 MB (fp8) ComfyUI/models/sam_audio/ audio_separate no-LLM mode AEmotionStudio/sam-audio
AudioX (Vocal Enhancement) ~1 GB ComfyUI/models/audiox/ AudioX chaining with ACE-Step AEmotionStudio/audiox
NormalCrafter (Surface Normals) ~2 GB ComfyUI/models/normalcrafter/ normalcrafter no-LLM mode AEmotionStudio/NormalCrafter
Kiwi-Edit (AI Video Editing) ~5 GB (FP8) / ~10 GB (BF16) ComfyUI/models/kiwi_edit_* kiwi_edit no-LLM mode AEmotionStudio/Kiwi-Edit-Instruct
DreamID-Omni (Talking Head) ⚠️ WIP ~12 GB (FP8) / ~23 GB (BF16) ComfyUI/models/dreamid_omni/ dreamid_omni no-LLM mode AEmotionStudio/dreamid-omni
FaceCam (Camera Control) ~16.8 GB (high+low bf16) ComfyUI/models/diffusion_models/ FaceCam node AEmotionStudio/facecam-wan2.2-14b-bf16
Fish Speech S2 Pro (TTS) ~6.5 GB (FP8) / ~10.4 GB (BF16) ComfyUI/models/fish_speech/ fish_speech no-LLM mode AEmotionStudio/fish-speech-s2-pro
Foundation-1 (Music Samples) ~2 GB ComfyUI/models/foundation1/ foundation1 no-LLM mode AEmotionStudio/foundation1-models
SeedVR2 (Diffusion Upscaler) ~3.4 GB (3B INT8) / ~8.3 GB (7B INT8) + ~300 MB (VAE) ComfyUI/models/SEEDVR2/ or ComfyUI/models/diffusion_models/ (both searched) ai_upscale with upscale_model = seedvr2_*. Prefer seedvr2_3b_int8 / seedvr2_7b_int8; raise blockswap_blocks to fit the 7B under 16 GB AEmotionStudio/SeedVR2-models (FP8/GGUF). INT8 ConvRot checkpoints are local-only β€” place seedvr2_{3b,7b}_int8_convrot.safetensors in either folder
FlashVSR v1.1 (One-Step VSR) ~5 GB (DiT + VAE + LQ_proj + TCDecoder) ComfyUI/models/FlashVSR/ ai_upscale with upscale_model = flashvsr_full / flashvsr_tiny / flashvsr_tiny_long AEmotionStudio/flashvsr-models β€” GPL-3.0
SCAIL-2 (Pose-Driven Animation) ~16 GB (fp8) ComfyUI/models/diffusion_models/ scail2 no-LLM mode. Also needs the Wan 2.1 VAE, UMT5-XXL text encoder and clip_vision_h Comfy-Org/SCAIL-2 β€” wan2.1_14B_SCAIL_2_fp8_scaled.safetensors
SVI 2.0 Pro (Infinite-Length Video) ~500 MB (high + low LoRAs) ComfyUI/models/svi/version-2.0/ svi no-LLM mode. ⚠️ Also requires the Wan 2.2 I2V-A14B base model (~28 GB) AEmotionStudio/svi-loras
Wan-Animate (Character Animation) ~18 GB (DiT) + ~2.5 GB (pose) + ~62 MB (detector) ComfyUI/models/wan_animate/ wan_animate no-LLM mode Kijai/WanVideo_comfy_fp8_scaled, Kijai/vitpose_comfy
MatAnyone2 (Video Matting) ~135 MB ComfyUI/models/matanyone2/ video_matting no-LLM mode AEmotionStudio/matanyone2 β€” ⚠️ NTU S-Lab 1.0, non-commercial
Sapiens2 (Human-Centric Vision) ~1 GB (0.4B) – ~10 GB (5B fp16) ComfyUI/models/sapiens2/<task>/ sapiens2 no-LLM mode β€” pose (308 keypoints), segmentation (29 classes), normals, pointmap, matting AEmotionStudio/sapiens2-normal and sibling sapiens2-* repos
SHARP (3D Gaussian View Synthesis) ~400 MB ComfyUI/models/sharp/ sharp no-LLM mode AEmotionStudio/sharp β€” ⚠️ Apple ML Research License, non-commercial
Visual Chronometer (PhyFPS) ~1 GB ComfyUI/models/visual_chronometer/ phyfps no-LLM mode AEmotionStudio/Visual_Chronometer

Note

Models are only downloaded when you use the corresponding skill for the first time. Core FFmpeg editing skills (200+ of them) require zero model downloads.


πŸŽ›οΈ Nodes

Every node, input and output β€” the Agent node's widgets, the save/load nodes, the Video Editor and the rest.

β†’ docs/NODES.md

🎯 Skill System

All 200+ skills by category, their parameters, and how to write your own.

β†’ docs/SKILLS.md

πŸ€– LLM Configuration

Ollama, the four CLI connectors (Gemini CLI, Claude Code, Cursor Agent, Qwen Code), model choice and troubleshooting.

β†’ docs/LLM-SETUP.md

πŸ› Troubleshooting

FFMPEG Not Found

Ensure FFMPEG is installed and in your system PATH:

ffmpeg -version

Install FFMPEG:

Platform Command / Method
Ubuntu / Debian sudo apt install ffmpeg
Arch / CachyOS sudo pacman -S ffmpeg
Fedora sudo dnf install ffmpeg
macOS brew install ffmpeg
Windows (winget) winget install Gyan.FFmpeg
Windows (choco) choco install ffmpeg
Windows (scoop) scoop install ffmpeg
Windows (manual) Download from ffmpeg.org/download, extract, and add the bin/ folder to your system PATH

Windows PATH tip: After installing, open a new terminal and run ffmpeg -version to verify. If not found, you may need to add ffmpeg's bin/ directory to your system PATH manually: Settings β†’ System β†’ About β†’ Advanced system settings β†’ Environment Variables β†’ Edit Path.

Ollama Connection Failed

Make sure Ollama is running:

ollama serve

If using a custom URL, set it in the node's ollama_url field.

Model Not Found

Pull the required model first:

ollama pull qwen2.5:8b
LLM Returns Empty Response

This usually means:

  • The model is still loading (first request after start)
  • The prompt is too long for the model's context window
  • Try running the same prompt again
  • Try a different model
Parameter Validation Errors

FFMPEGA auto-coerces types (float→int) and clamps out-of-range values. If you still see errors, try simplifying your prompt or using a more capable model.

Cancelling a Running Request

If the LLM is taking too long or you want to abort mid-request, close the ComfyUI terminal or restart ComfyUI instead of using the interrupt button. The interrupt button waits for the current LLM response to complete, which can take a while β€” closing/restarting ComfyUI kills it immediately.


🀝 Contributing

Contributions are welcome! Whether it's bug reports, new skills, or improvements β€” your help is appreciated.

  1. Fork the Project
  2. Create your Feature Branch (git checkout -b feature/AmazingFeature)
  3. Commit your Changes (git commit -m 'Add some AmazingFeature')
  4. Push to the Branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

πŸ“ License

This project is licensed under the GPL-3.0 License β€” see the LICENSE file for details.


Developed by Γ†motion Studio

YouTube Discord Ko-fi

Releases

Packages

Contributors

Languages