Reproducible case study of pitfalls in contrastive SAE discovery and steering for "consciousness" features (GemmaScope SAEs, Gemma 3 4B/12B): reconstruction confound, delta-steering fix, matched controls, and false-positive scaling law vs dataset size.
gemma sae sparse-autoencoder contrastive-learning mechanistic-interpretability feature-steering neuronpedia null-result gemmascope delta-steering
-
Updated
Feb 26, 2026 - Python