BenchMIRT:一种新的LLM基准审计方法
BenchMIRT 是一种新的方法,用于在单个提示级别审计大型语言模型(LLM)基准,揭示现有基准的见解,并建议更有效的评估方法。
BenchMIRT 是一种新的方法,用于在单个提示级别审计大型语言模型(LLM)基准,揭示现有基准的见解,并建议更有效的评估方法。
Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.
BenchMIRT applies IRT at both the model and question level. For a given model, it estimates the model’s strength on the capabilities reflected across the selected benchmarks.