本文へ移動
安定報道日時 2026-10-02 12:00

制御されたワールドモデル適応における校正リスクルーティング

研究者たちは、モデルベースの強化学習において新しいアプローチであるModel-Corrected World Model (MC-WM)を紹介しました。これはシミュレータにおけるモデル選択問題を解決するために、ターゲットデータを分割し、信頼信号を使用してモデルを適応的に調整する方法です。

01

根拠

  • AarXiv cs.AI一次情報2026-10-02 12:00
    Abstract: Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions a…
    出典を見る