Worked Gemma Scope case study

Design of Experiments for Sparse Autoencoder Feature Stability

Abstract

Sparse autoencoders offer a practical way to transform transformer residual-stream activations into sparse feature vectors, but it is not yet clear how stable those feature vectors are under controlled prompt perturbations. This report frames SAE activations as response variables in a classical Design of Experiments setting. Using 405 unique prompts crossing content domain, role framing, surface paraphrase, and language register at three transformer layers, the study shows that layer dominates the measured responses while substantial layer-by-content and layer-by-role interactions reorder the apparent importance of prompt factors. The contribution is a measurement scaffold for deciding which factors and interactions should be interpreted before undertaking causal circuit analysis.