Official implementation and resources — coming soon.
Note
This repository is currently under preparation. Code will be released here.
ShieldCLIP is a selective safety-alignment framework for CLIP-like multimodal encoders. It conditions alignment on the observed safety state of each modality, preserving benign representations while redirecting only unsafe content.
The framework is trained with ViSUv2, a dataset of 195K real/generated image-text quadruplets with independent safety labels for text and images across 578 concepts and 28 harmful-content categories.