ActiveSaddler:用于代理框架优化的自动化课程学习
一种名为ActiveSaddler的新方法通过根据 harness 的不断变化的需求调整训练课程,实现了 LLM 代理 harness 的自动化优化。
一种名为ActiveSaddler的新方法通过根据 harness 的不断变化的需求调整训练课程,实现了 LLM 代理 harness 的自动化优化。
We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler.
ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets.
However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as…