跳到正文
Berkeley AI Research·· 2026-04-20

Gradient-based Planning for World Models at Longer Horizons

Gradient-based Planning for World Models at Longer Horizons

中文摘要

GRASP 是一种基于梯度的新型规划器,用于学习到的动力学(“世界模型”),它通过以下方式使长期规划变得实用:(1) 将轨迹提升到虚拟状态,从而使优化在时间上并行进行;(2) 直接向状态迭代中添加随机性以进行探索;(3) 重新塑造梯度,使动作获得清晰的信号,同时通过高维视觉模型避免脆弱的“状态-输入”梯度。大型的学习世界模型正变得越来越强大。它们可以在高维视觉空间中预测长时间序列的未来观察结果,并以几年前难以想象的方式在不同任务之间进行泛化。随着这些模型的扩展,它们开始看起来不像特定任务的预测器,而更像通用的模拟器。但拥有一个强大的预测模型并不等于能够有效地将其用于控制/学习/规划。实际上,使用现代世界模型进行长期规划仍然很脆弱:优化变得条件不良,非贪心结构会产生不良的局部最小值,而高维潜在空间则引入了微妙的故障模式。在这篇博客文章中,我将描述促使这项研究的问题。

英文原文

GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models. Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators. But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes. In this blog post, I describe the problems that motivated this pro

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:Berkeley AI Research · bair.berkeley.edu