Zhuochun Li demonstrates that small language models encode evaluative signals in their hidden states despite weak generative ability. The work proposes a "Representation-as-a-Judge" paradigm using INSPECTOR, a probing-based framework that predicts aspect-level evaluation scores from small model representations without decoding. Experiments on GSM8K, MATH, and GPQA show INSPECTOR outperforms prompting-based small LMs and closely approximates full LLM judges, offering a more efficient and interpretable alternative for scalable evaluation.
No score is assigned. Sources and their independence are shown in the citation chain below.