数学的推論能力を価値に基づいた情報検索を通じて向上させる
新しいトレーニングフレームワークであるAPIVISは、直接的な応答と検索された応答を組み合わせることで、大規模言語モデルの数学的推論能力を向上させ、検証可能な報酬を用いた強化学習を活用しています。
新しいトレーニングフレームワークであるAPIVISは、直接的な応答と検索された応答を組み合わせることで、大規模言語モデルの数学的推論能力を向上させ、検証可能な報酬を用いた強化学習を活用しています。
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models.
we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning
Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training…