ResearchHugging Face

BenchMIRT: What are LLM benchmarks actually measuring?

#llm#benchmarking#safety#reasoning#ai

English

BenchMIRT is a new method for auditing LLM benchmarks, analyzing individual prompts to reveal the underlying capabilities measured by benchmarks. It uses multidimensional Item Response Theory to separate signals related to safety and general reasoning, highlighting complexities in existing benchmarks like BBQ and WMDP.

中文

BenchMIRT是一种新的LLM基准审计方法,通过分析单个提示揭示基准测量的潜在能力。它使用多维项目反应理论来分离与安全性和一般推理相关的信号,突显了现有基准(如BBQ和WMDP)的复杂性。