跳到正文
每分钟自动更新
11月1日周六
  1. Berkeley AI Research

    RL without TD learning

    中文摘要

    在这篇文章中,我将介绍一种基于“分而治之”(divide and conquer)这一“替代”范式的强化学习(RL)算法。与传统方法不同,该算法并不基于时间差分(TD)学习(这存在可扩展性问题),并且能够很好地扩展到长时域任务中。我们可以基于分而治之的方法进行强化学习(RL),而不是基于时间差分(TD)学习。问题设定:离策略RL。我们的问题设定是离策略RL。让我们简要回顾一下这意味着什么。在强化学习中,有两种算法类别:在策略(on-policy)RL和离策略(off-policy)RL。在策略RL意味着我们只能使用当前策略收集的新数据。换句话说,每次更新策略时,我们都必须丢弃旧数据。PPO和GRPO(以及一般的策略梯度方法)都属于这一类。离策略RL意味着我们没有这种限制:我们可以使用任何类型的数据,包括旧的经验、人类示范、互联网数据等等。因此,离策略RL比在策略RL更通用和灵活(当然也更困难!)。Q-learning是最著名的离策略RL算法。在数据收集成本较高的领域(例如)

    英文原文

    In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks. We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning. Problem setting: off-policy RL Our problem setting is off-policy RL . Let’s briefly review what this means. There are two classes of algorithms in RL: on-policy RL and off-policy RL. On-policy RL means we can only use fresh data collected by the current policy. In other words, we have to throw away old data each time we update the policy. Algorithms like PPO and GRPO (and policy gradient methods in general) belong to this category. Off-policy RL means we don’t have this restriction: we can use any kind of data, including old experience, human demonstrations, Internet data, and so on. So off-policy RL is more general and flexible than on-policy RL (and of course harder!). Q-learning is the most well-known off-policy RL algorithm. In domains where data collection is expensive ( e.g.

9月1日周一
  1. Berkeley AI Research

    What exactly does word2vec learn?

    中文摘要

    word2vec到底学到了什么,又是如何学习的呢?回答这个问题等同于在一种最小但有趣的语言建模任务中理解表示学习。尽管word2vec是现代语言模型的一个著名先驱,但多年来,研究人员缺乏一个定量且具有预测性的理论来描述其学习过程。在我们最新的论文中,我们终于提供了这样的理论。我们证明,在一些现实且实用的条件下,学习问题可以简化为无权重的最小二乘矩阵分解。我们以闭合形式求解了梯度流动力学;最终学到的表示形式仅仅是通过PCA得到的。word2vec的学习动态。当从较小的初始化开始训练时,word2vec以离散的、顺序的步骤进行学习。左图:权重矩阵中的秩增加学习步骤,每个步骤都减少损失。右图:三个时间切片的潜在嵌入空间,展示了嵌入向量如何在每个学习步骤中扩展为维度不断增加的子空间,直到模型容量达到饱和。在详细阐述这一结果之前,让我们先说明这个问题。word2vec是一个著名的用于学习密集向量表示的算法。

    英文原文

    What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known precursor to modern language models, for many years, researchers lacked a quantitative and predictive theory describing its learning process. In our new paper , we finally provide such a theory. We prove that there are realistic, practical regimes in which the learning problem reduces to unweighted least-squares matrix factorization . We solve the gradient flow dynamics in closed form; the final learned representations are simply given by PCA. Learning dynamics of word2vec . When trained from small initialization, word2vec learns in discrete, sequential steps. Left: rank-incrementing learning steps in the weight matrix, each decreasing the loss. Right: three time slices of the latent embedding space showing how embedding vectors expand into subspaces of increasing dimension at each learning step, continuing until model capacity is saturated. Before elaborating on this result, let’s motivate the problem. word2vec is a well-known algorithm for learning dense vector repres