What exactly does word2vec learn?
What exactly does word2vec learn?
word2vec到底学到了什么,又是如何学习的呢?回答这个问题等同于在一种最小但有趣的语言建模任务中理解表示学习。尽管word2vec是现代语言模型的一个著名先驱,但多年来,研究人员缺乏一个定量且具有预测性的理论来描述其学习过程。在我们最新的论文中,我们终于提供了这样的理论。我们证明,在一些现实且实用的条件下,学习问题可以简化为无权重的最小二乘矩阵分解。我们以闭合形式求解了梯度流动力学;最终学到的表示形式仅仅是通过PCA得到的。word2vec的学习动态。当从较小的初始化开始训练时,word2vec以离散的、顺序的步骤进行学习。左图:权重矩阵中的秩增加学习步骤,每个步骤都减少损失。右图:三个时间切片的潜在嵌入空间,展示了嵌入向量如何在每个学习步骤中扩展为维度不断增加的子空间,直到模型容量达到饱和。在详细阐述这一结果之前,让我们先说明这个问题。word2vec是一个著名的用于学习密集向量表示的算法。
What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known precursor to modern language models, for many years, researchers lacked a quantitative and predictive theory describing its learning process. In our new paper , we finally provide such a theory. We prove that there are realistic, practical regimes in which the learning problem reduces to unweighted least-squares matrix factorization . We solve the gradient flow dynamics in closed form; the final learned representations are simply given by PCA. Learning dynamics of word2vec . When trained from small initialization, word2vec learns in discrete, sequential steps. Left: rank-incrementing learning steps in the weight matrix, each decreasing the loss. Right: three time slices of the latent embedding space showing how embedding vectors expand into subspaces of increasing dimension at each learning step, continuing until model capacity is saturated. Before elaborating on this result, let’s motivate the problem. word2vec is a well-known algorithm for learning dense vector repres
应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。
来源:Berkeley AI Research · bair.berkeley.edu