Mark Oskin's research shows that a transformer built from legible, bounded operators requires a per-channel variance floor during training to prevent a crispness penalty from collapsing detectors into dead constants. The resulting model routes 87% of load-bearing computation through crisp operators, achieving 78% legible feed-forward operands while maintaining quality parity with a conventional baseline.
No score is assigned. Sources and their independence are shown in the citation chain below.