TLDR;

In previous work, we found a problematic form a feature splitting called "feature absorption" when analyzing Gemma Scope SAEs. We hypothesized that this was due to SAEs struggling to separate co-occurrence between features, but we did not prove this. In this post, we set up toy models where we can explicitly control feature representations and co-occurrence rates and show the following:

Feature absorption happens when features co-occur.
If co-occurring feature magnitudes vary relative to each other, we observe "partial absorption", where a latent tracking a main feature sometimes fires weakly instead of not firing at all, but sometimes does fully not fire.
Feature absorption happens even with imperfect co-occurrence, depending on the strength of the sparsity penalty.
Tying the SAE encoder and decoder weights together solves feature absorption.