BenchMIRT: LLMベンチマークの監査に新しい手法
BenchMIRTは、大規模言語モデル(LLM)ベンチマークを個々のプロンプトレベルで監査する新しい手法であり、既存のベンチマークに関する洞察を明らかにし、より効率的な評価方法を提案します。
BenchMIRTは、大規模言語モデル(LLM)ベンチマークを個々のプロンプトレベルで監査する新しい手法であり、既存のベンチマークに関する洞察を明らかにし、より効率的な評価方法を提案します。
Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.
BenchMIRT applies IRT at both the model and question level. For a given model, it estimates the model’s strength on the capabilities reflected across the selected benchmarks.