跳到正文
稳定技术报道时间 2026-09-02 05:39

BenchMIRT:一种新的LLM基准审计方法

BenchMIRT 是一种新的方法,用于在单个提示级别审计大型语言模型(LLM)基准,揭示现有基准的见解,并建议更有效的评估方法。

01

证据

  • HHugging Face Blog公司披露2026-09-02 05:39
    Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on.
    查看来源
  • HHugging Face Blog公司披露2026-09-02 05:39
    BenchMIRT applies IRT at both the model and question level. For a given model, it estimates the model’s strength on the capabilities reflected across the selected benchmarks.
    查看来源