Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

ShieldCLIP

Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

Official implementation and resources — coming soon.

Note

This repository is currently under preparation. Code will be released here.

Overview

ShieldCLIP is a selective safety-alignment framework for CLIP-like multimodal encoders. It conditions alignment on the observed safety state of each modality, preserving benign representations while redirecting only unsafe content.

The framework is trained with ViSUv2, a dataset of 195K real/generated image-text quadruplets with independent safety labels for text and images across 578 concepts and 28 harmful-content categories.

About

Official repository for ShieldCLIP: selective safety alignment for harmful content mitigation in multimodal foundation models. Code, trained models, and dataset access instructions coming soon.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors