Multiverse's paper on LLM compression reports that removing 50% of Llama-3.3-70B-Instruct's blocks yields almost 23 percentage points more on MMLU than the best competing block-removal method. The approach reformulates block selection as a constrained binary optimization problem mapped to an Ising glass, using a Hessian computed once from calibration data as a proxy for benchmark quality, solved with classical and quantum-inspired solvers.
No score is assigned. Sources and their independence are shown in the citation chain below.