RISED: Agentic Multi-environment Selection and Self-DistillationのためのRubrics
新しい方法RISEDが導入され、多様なインタラクティブ環境におけるLLMエージェントのトレーニングに焦点を当てています。これはプロンプトグループの選択に加え、環境間での学習率の変化を処理することを目的としています。
新しい方法RISEDが導入され、多様なインタラクティブ環境におけるLLMエージェントのトレーニングに焦点を当てています。これはプロンプトグループの選択に加え、環境間での学習率の変化を処理することを目的としています。
An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments.
Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment.
Abstract: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-gro…