跳到正文
Berkeley AI Research·· 2025-11-01

RL without TD learning

RL without TD learning

中文摘要

在这篇文章中,我将介绍一种基于“分而治之”(divide and conquer)这一“替代”范式的强化学习(RL)算法。与传统方法不同,该算法并不基于时间差分(TD)学习(这存在可扩展性问题),并且能够很好地扩展到长时域任务中。我们可以基于分而治之的方法进行强化学习(RL),而不是基于时间差分(TD)学习。问题设定:离策略RL。我们的问题设定是离策略RL。让我们简要回顾一下这意味着什么。在强化学习中,有两种算法类别:在策略(on-policy)RL和离策略(off-policy)RL。在策略RL意味着我们只能使用当前策略收集的新数据。换句话说,每次更新策略时,我们都必须丢弃旧数据。PPO和GRPO(以及一般的策略梯度方法)都属于这一类。离策略RL意味着我们没有这种限制:我们可以使用任何类型的数据,包括旧的经验、人类示范、互联网数据等等。因此,离策略RL比在策略RL更通用和灵活(当然也更困难!)。Q-learning是最著名的离策略RL算法。在数据收集成本较高的领域(例如)

英文原文

In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks. We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning. Problem setting: off-policy RL Our problem setting is off-policy RL . Let’s briefly review what this means. There are two classes of algorithms in RL: on-policy RL and off-policy RL. On-policy RL means we can only use fresh data collected by the current policy. In other words, we have to throw away old data each time we update the policy. Algorithms like PPO and GRPO (and policy gradient methods in general) belong to this category. Off-policy RL means we don’t have this restriction: we can use any kind of data, including old experience, human demonstrations, Internet data, and so on. So off-policy RL is more general and flexible than on-policy RL (and of course harder!). Q-learning is the most well-known off-policy RL algorithm. In domains where data collection is expensive ( e.g.

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:Berkeley AI Research · bair.berkeley.edu