跳到正文
稳定技术报道时间 2026-10-02 12:00

基于依赖关系的奖励塑造用于代理强化学习

研究人员提出依赖感知奖励塑造(DARS),通过根据先决条件关系分配步骤级信用来改进大型语言模型中的强化学习,解决失败情节中浪费努力的问题。

01

证据

  • AarXiv cs.AI一手来源2026-10-02 12:00
    A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers.
    查看来源
  • AarXiv cs.AI一手来源2026-10-02 12:00
    We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph.
    查看来源
  • AarXiv cs.AI一手来源2026-10-02 12:00
    Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it mad…
    查看来源