Mechanistic study of the refusal direction across base, instruction-tuned, and reasoning-distilled Qwen2.5-1.5B variants: extraction, ablation, transplant, and phase-aware analysis.
-
Updated
May 8, 2026 - Python
Mechanistic study of the refusal direction across base, instruction-tuned, and reasoning-distilled Qwen2.5-1.5B variants: extraction, ablation, transplant, and phase-aware analysis.
Post-hoc refusal ablation ("abliteration") for qwen3_5 hybrid-attention models — the architecture TransformerLens-based tools cannot load. Training-free weight edit: refusal 78.1% -> 9.4%, capability unchanged.
Studying how safety alignment is encoded in LLM weights, using Arbitrary-Rank Ablation to isolate and remove the refusal direction in Gemma 4 E2B.
Carrier-probe benchmark that measures the 'flinch' — how much a model shrinks the probability of a charged word when it is the obvious next token in a sentence.
Statistical evaluation of Gemma refusal robustness using direction ablation and official SORRY-Bench scoring.
To associate your repository with the refusal-ablation topic, visit your repo's landing page and select "manage topics."