本文へ移動
安定テクノロジー報道日時 2026-10-02 12:00

数学的推論能力を価値に基づいた情報検索を通じて向上させる

新しいトレーニングフレームワークであるAPIVISは、直接的な応答と検索された応答を組み合わせることで、大規模言語モデルの数学的推論能力を向上させ、検証可能な報酬を用いた強化学習を活用しています。

01

誰に影響するか

  1. 1APIVIS
  2. と競合 →事実
  3. を使用 →事実
  4. に依存 →事実
    4GRPO技術
  5. を可能にする →事実
  6. を必要とする →事実
02

根拠

  • AarXiv cs.AI一次情報2026-10-02 12:00
    Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models.
    出典を見る
  • AarXiv cs.AI一次情報2026-10-02 12:00
    we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning
    出典を見る
  • AarXiv cs.AI一次情報2026-10-02 12:00
    Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
    出典を見る
  • AarXiv cs.AI一次情報2026-10-02 12:00
    Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training…
    出典を見る