Skip to content
StableTechnologyReported 2026-10-02 12:00

Improving Math Reasoning through Value-guided Informative Search

A new training framework, APIVIS, combines direct and searched responses to enhance the mathematical reasoning capabilities of large language models using reinforcement learning with verifiable rewards.

01

Who it touches

  1. 1APIVIS
  2. competes with →Fact
  3. uses →Fact
  4. depends on →Fact
    4GRPOTechnology
  5. enables →Fact
  6. requires →Fact
02

Evidence

  • AarXiv cs.AIPrimary source2026-10-02 12:00
    Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models.
    View source
  • AarXiv cs.AIPrimary source2026-10-02 12:00
    we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning
    View source
  • AarXiv cs.AIPrimary source2026-10-02 12:00
    Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
    View source
  • AarXiv cs.AIPrimary source2026-10-02 12:00
    Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training…
    View source