Improving Math Reasoning through Value-guided Informative Search
A new training framework, APIVIS, combines direct and searched responses to enhance the mathematical reasoning capabilities of large language models using reinforcement learning with verifiable rewards.