Moritz Miller proposes constraining internal features in language models to be almost orthogonal to reduce feature entanglement and enable more isolated causal interventions. The approach formalizes feature interference, upper-bounds its propagation in terms of self-coherence, and connects this to an explicit orthogonality regularization on the feature dictionary. Empirical results show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving overall model performance.
No score is assigned. Sources and their independence are shown in the citation chain below.