Meera Desai and colleagues analyzed validation practices across papers from eight flagship social science journals that use large language models as measurement instruments. The study finds that LLM-generated measurements frequently serve a central role in empirical analyses, but validation practices remain "inconsistent and limited." The researchers outline complementary strategies for more robust validation, aiming to establish better norms and standards for addressing known challenges such as bias, hallucination, and brittleness when using LLMs in social science research.
No score is assigned. Sources and their independence are shown in the citation chain below.