稳定报道时间 2026-10-02 12:00
受控世界模型适应的校准风险路由
研究人员介绍了模型校正的世界模型(MC-WM),这是一种新的基于模型的强化学习方法,通过划分目标数据并使用置信度信号来调整模型,以解决模拟器中的模型选择问题。
研究人员介绍了模型校正的世界模型(MC-WM),这是一种新的基于模型的强化学习方法,通过划分目标数据并使用置信度信号来调整模型,以解决模拟器中的模型选择问题。
Abstract: Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions a…