BenchMIRT: A New Method for Auditing LLM Benchmarks
BenchMIRT is a new method for auditing large language model (LLM) benchmarks at the level of individual prompts, revealing insights into existing benchmarks and suggesting more efficient evaluation methods.