← Back to the wire

BenchMIRT: What are LLM benchmarks actually measuring?

AnnouncementResearchSep 1, 2026

Researchers introduced BenchMIRT, a method using multidimensional Item Response Theory to audit LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks and over 34,000 questions, it independently recovered two stable dimensions: safety and general reasoning. The analysis found some benchmarks mix signals—for example, BBQ, whose questions include an Uber booking scenario probing age bias, aligned more with general reasoning than safety, as did WMDP and HarmBench's copyright questions.

Receipt № 16831 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
UberCompany
Canonical: https://huggingface.co/blog/allenai/benchmirt