跳到正文
每分钟自动更新
5月23日周六
5月18日周一
5月11日周一
5月8日周五
  1. Berkeley AI Research

    Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

    中文摘要

    自适应并行推理概述。如果一个推理模型可以自行决定何时分解和并行化独立的子任务,生成多少个并发线程,以及如何根据具体问题协调这些线程,会怎样呢?我们对并行推理领域的最新进展进行了详细分析,特别是自适应并行推理。免责声明:本文一部分是对现状的综述,另一部分是对自适应并行推理的见解。其中一位作者(Tony Lian)共同领导了ThreadWeaver(Lian等人,2025年),这是下文讨论的方法之一。作者的目标是根据每种方法自身的特性来呈现它们。动机 近年来,大语言模型推理能力的进展主要由推理时的扩展驱动,除了数据和参数的扩展(OpenAI等,2024;DeepSeek-AI等,2025)。那些显式输出推理标记(通过中间步骤、回溯和探索)的模型现在在数学、编程和代理基准测试中占据主导地位。这些行为使模型能够探索替代假设,纠正早期的错误,并综合得出结论,而不是坚持单一解决方案(Wen等,2025)。问题在于序列

    英文原文

    Overview of adaptive parallel reasoning. What if a reasoning model could decide for itself when to decompose and parallelize independent subtasks, how many concurrent threads to spawn, and how to coordinate them based on the problem at hand? We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning. Disclosure: this post is part landscape survey, part perspective on adaptive parallel reasoning. One of the authors (Tony Lian) co-led ThreadWeaver ( Lian et al., 2025 ), one of the methods discussed below. The authors aim to present each approach on its own terms. Motivation Recent progress in LLM reasoning capabilities has been largely driven by inference-time scaling, in addition to data and parameter scaling ( OpenAI et al., 2024 ; DeepSeek-AI et al., 2025 ). Models that explicitly output reasoning tokens (through intermediate steps, backtracking, and exploration) now dominate math, coding, and agentic benchmarks. These behaviors allow models to explore alternative hypotheses, correct earlier mistakes, and synthesize conclusions rather than committing to a single solution ( Wen et al., 2025 ). The problem is that seq

4月20日周一
  1. Berkeley AI Research

    Gradient-based Planning for World Models at Longer Horizons

    中文摘要

    GRASP 是一种基于梯度的新型规划器,用于学习到的动力学(“世界模型”),它通过以下方式使长期规划变得实用:(1) 将轨迹提升到虚拟状态,从而使优化在时间上并行进行;(2) 直接向状态迭代中添加随机性以进行探索;(3) 重新塑造梯度,使动作获得清晰的信号,同时通过高维视觉模型避免脆弱的“状态-输入”梯度。大型的学习世界模型正变得越来越强大。它们可以在高维视觉空间中预测长时间序列的未来观察结果,并以几年前难以想象的方式在不同任务之间进行泛化。随着这些模型的扩展,它们开始看起来不像特定任务的预测器,而更像通用的模拟器。但拥有一个强大的预测模型并不等于能够有效地将其用于控制/学习/规划。实际上,使用现代世界模型进行长期规划仍然很脆弱:优化变得条件不良,非贪心结构会产生不良的局部最小值,而高维潜在空间则引入了微妙的故障模式。在这篇博客文章中,我将描述促使这项研究的问题。

    英文原文

    GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models. Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators. But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes. In this blog post, I describe the problems that motivated this pro

3月13日周五
  1. Berkeley AI Research

    Identifying Interactions at Scale for LLMs

    中文摘要

    理解复杂机器学习系统的行为,尤其是大型语言模型(LLMs),是现代人工智能中的一个关键挑战。可解释性研究旨在使决策过程对模型开发者和受影响的人类更加透明,这是迈向更安全、更可信赖的人工智能的重要一步。为了全面理解这些系统,我们可以从不同的角度进行分析:特征归因(feature attribution),它隔离出推动预测的具体输入特征(Lundberg & Lee, 2017;Ribeiro 等人,2022);数据归因(data attribution),它将模型行为与有影响力的训练样本联系起来(Koh & Liang, 2017;Ilyas 等人,2022);以及机制可解释性(mechanistic interpretability),它剖析内部组件的功能(Conmy 等人,2023;Sharkey 等人,2025)。在这些视角中,同样的基本障碍依然存在:规模上的复杂性。模型行为很少是孤立组件的结果;相反,它源于复杂的依赖关系和模式。为了实现最先进的性能,模型会综合复杂的特征关系,从多样化的训练样本中寻找共同模式,并通过处理信息来实现。

    英文原文

    Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decision-making process more transparent to model builders and impacted humans, a step toward safer and more trustworthy AI. To gain a comprehensive understanding, we can analyze these systems through different lenses: feature attribution , which isolates the specific input features driving a prediction ( Lundberg & Lee, 2017 ; Ribeiro et al., 2022 ); data attribution , which links model behaviors to influential training examples ( Koh & Liang, 2017 ; Ilyas et al., 2022 ); and mechanistic interpretability , which dissects the functions of internal components ( Conmy et al., 2023 ; Sharkey et al., 2025 ). Across these perspectives, the same fundamental hurdle persists: complexity at scale . Model behavior is rarely the result of isolated components; rather, it emerges from complex dependencies and patterns. To achieve state-of-the-art performance, models synthesize complex feature relationships, find shared patterns from diverse training examples, and process information throug

1月10日周六
  1. Berkeley AI Research

    Information-Driven Design of Imaging Systems

    中文摘要

    编码器(光学系统)将物体映射到无噪声的图像,而噪声会将这些图像转化为测量值。我们的信息估计器仅使用这些带有噪声的测量值和噪声模型,来量化测量值区分物体的能力。许多成像系统产生的测量值人类根本无法看到,或者无法直接理解。你的智能手机在生成最终照片之前,会通过算法处理原始传感器数据。磁共振成像(MRI)扫描仪收集的是频率空间中的测量值,这些测量值需要经过重建后,医生才能查看。自动驾驶汽车会直接通过神经网络处理摄像头和激光雷达(LiDAR)数据。在这些系统中,重要的不是测量值看起来是什么样子,而是它们包含多少有用的信息。即使信息以人类无法理解的方式编码,人工智能也能提取这些信息。然而,我们很少直接评估信息内容。传统的指标如分辨率和信噪比分别评估质量的各个单独方面,这使得在这些因素之间存在权衡的系统之间进行比较变得困难。常见的替代方法是训练神经网络来重建或分类图像,这将成像硬件的质量与神经网络的质量混为一谈。

    英文原文

    An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements distinguish objects. Many imaging systems produce measurements that humans never see or cannot interpret directly. Your smartphone processes raw sensor data through algorithms before producing the final photo. MRI scanners collect frequency-space measurements that require reconstruction before doctors can view them. Self-driving cars process camera and LiDAR data directly with neural networks. What matters in these systems is not how measurements look, but how much useful information they contain. AI can extract this information even when it is encoded in ways that humans cannot interpret. And yet we rarely evaluate information content directly. Traditional metrics like resolution and signal-to-noise ratio assess individual aspects of quality separately, making it difficult to compare systems that trade off between these factors. The common alternative, training neural networks to reconstruct or classify images, conflates the quality of the imaging hardware with the qual

11月1日周六
  1. Berkeley AI Research

    RL without TD learning

    中文摘要

    在这篇文章中,我将介绍一种基于“分而治之”(divide and conquer)这一“替代”范式的强化学习(RL)算法。与传统方法不同,该算法并不基于时间差分(TD)学习(这存在可扩展性问题),并且能够很好地扩展到长时域任务中。我们可以基于分而治之的方法进行强化学习(RL),而不是基于时间差分(TD)学习。问题设定:离策略RL。我们的问题设定是离策略RL。让我们简要回顾一下这意味着什么。在强化学习中,有两种算法类别:在策略(on-policy)RL和离策略(off-policy)RL。在策略RL意味着我们只能使用当前策略收集的新数据。换句话说,每次更新策略时,我们都必须丢弃旧数据。PPO和GRPO(以及一般的策略梯度方法)都属于这一类。离策略RL意味着我们没有这种限制:我们可以使用任何类型的数据,包括旧的经验、人类示范、互联网数据等等。因此,离策略RL比在策略RL更通用和灵活(当然也更困难!)。Q-learning是最著名的离策略RL算法。在数据收集成本较高的领域(例如)

    英文原文

    In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks. We can do Reinforcement Learning (RL) based on divide and conquer, instead of temporal difference (TD) learning. Problem setting: off-policy RL Our problem setting is off-policy RL . Let’s briefly review what this means. There are two classes of algorithms in RL: on-policy RL and off-policy RL. On-policy RL means we can only use fresh data collected by the current policy. In other words, we have to throw away old data each time we update the policy. Algorithms like PPO and GRPO (and policy gradient methods in general) belong to this category. Off-policy RL means we don’t have this restriction: we can use any kind of data, including old experience, human demonstrations, Internet data, and so on. So off-policy RL is more general and flexible than on-policy RL (and of course harder!). Q-learning is the most well-known off-policy RL algorithm. In domains where data collection is expensive ( e.g.

9月1日周一
  1. Berkeley AI Research

    What exactly does word2vec learn?

    中文摘要

    word2vec到底学到了什么,又是如何学习的呢?回答这个问题等同于在一种最小但有趣的语言建模任务中理解表示学习。尽管word2vec是现代语言模型的一个著名先驱,但多年来,研究人员缺乏一个定量且具有预测性的理论来描述其学习过程。在我们最新的论文中,我们终于提供了这样的理论。我们证明,在一些现实且实用的条件下,学习问题可以简化为无权重的最小二乘矩阵分解。我们以闭合形式求解了梯度流动力学;最终学到的表示形式仅仅是通过PCA得到的。word2vec的学习动态。当从较小的初始化开始训练时,word2vec以离散的、顺序的步骤进行学习。左图:权重矩阵中的秩增加学习步骤,每个步骤都减少损失。右图:三个时间切片的潜在嵌入空间,展示了嵌入向量如何在每个学习步骤中扩展为维度不断增加的子空间,直到模型容量达到饱和。在详细阐述这一结果之前,让我们先说明这个问题。word2vec是一个著名的用于学习密集向量表示的算法。

    英文原文

    What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known precursor to modern language models, for many years, researchers lacked a quantitative and predictive theory describing its learning process. In our new paper , we finally provide such a theory. We prove that there are realistic, practical regimes in which the learning problem reduces to unweighted least-squares matrix factorization . We solve the gradient flow dynamics in closed form; the final learned representations are simply given by PCA. Learning dynamics of word2vec . When trained from small initialization, word2vec learns in discrete, sequential steps. Left: rank-incrementing learning steps in the weight matrix, each decreasing the loss. Right: three time slices of the latent embedding space showing how embedding vectors expand into subspaces of increasing dimension at each learning step, continuing until model capacity is saturated. Before elaborating on this result, let’s motivate the problem. word2vec is a well-known algorithm for learning dense vector repres