跳到正文
稳定技术报道时间 2026-10-02 12:00

通过价值引导的信息搜索提高数学推理能力

一种新的训练框架APIVIS结合直接响应和搜索响应,通过可验证奖励的强化学习来增强大型语言模型的数学推理能力。

01

影响到谁

  1. 1APIVIS
  2. 竞争 →事实
  3. 使用 →事实
  4. 依赖 →事实
    4GRPO技术
  5. 支撑 →事实
  6. 需要 →事实
02

证据

  • AarXiv cs.AI一手来源2026-10-02 12:00
    Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models.
    查看来源
  • AarXiv cs.AI一手来源2026-10-02 12:00
    we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning
    查看来源
  • AarXiv cs.AI一手来源2026-10-02 12:00
    Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
    查看来源
  • AarXiv cs.AI一手来源2026-10-02 12:00
    Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training…
    查看来源