本文へ移動
安定テクノロジー報道日時 2026-10-02 12:00

Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

研究者たちは、依存関係に基づいた報酬形成(Dependency-Aware Reward Shaping、DARS)を提案し、大規模言語モデルにおける強化学習を改善することを目的としています。このアプローチでは、前提条件の関係に基づいてステップレベルのクレジットを割り当てることで、失敗したエピソードにおける無駄な努力の問題に対処します。

01

根拠

  • AarXiv cs.AI一次情報2026-10-02 12:00
    A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers.
    出典を見る
  • AarXiv cs.AI一次情報2026-10-02 12:00
    We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph.
    出典を見る
  • AarXiv cs.AI一次情報2026-10-02 12:00
    Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it mad…
    出典を見る