ActiveSaddler: エージェントハーゴス最適化のための自動カリキュラム学習
新しい方法であるActiveSaddlerは、LLMエージェントのハーキスを自動的に最適化することにより、ハーキスの進化するニーズに応じてトレーニングカリキュラムを調整します。
新しい方法であるActiveSaddlerは、LLMエージェントのハーキスを自動的に最適化することにより、ハーキスの進化するニーズに応じてトレーニングカリキュラムを調整します。
We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler.
ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets.
However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as…