Reproducible toy-model experiments in mechanistic interpretability. Includes a test of whether mixed (superposed) coding is what lets a linear readout see feature interactions — headline result plus a reported null.
pytorch computational-neuroscience superposition ai-safety interpretability toy-models neuroai mechanistic-interpretability polysemanticity mixed-selectivity
-
Updated
Jul 27, 2026 - Python