跳到正文
Berkeley AI Research·· 2026-05-08

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling

中文摘要

自适应并行推理概述。如果一个推理模型可以自行决定何时分解和并行化独立的子任务,生成多少个并发线程,以及如何根据具体问题协调这些线程,会怎样呢?我们对并行推理领域的最新进展进行了详细分析,特别是自适应并行推理。免责声明:本文一部分是对现状的综述,另一部分是对自适应并行推理的见解。其中一位作者(Tony Lian)共同领导了ThreadWeaver(Lian等人,2025年),这是下文讨论的方法之一。作者的目标是根据每种方法自身的特性来呈现它们。动机 近年来,大语言模型推理能力的进展主要由推理时的扩展驱动,除了数据和参数的扩展(OpenAI等,2024;DeepSeek-AI等,2025)。那些显式输出推理标记(通过中间步骤、回溯和探索)的模型现在在数学、编程和代理基准测试中占据主导地位。这些行为使模型能够探索替代假设,纠正早期的错误,并综合得出结论,而不是坚持单一解决方案(Wen等,2025)。问题在于序列

英文原文

Overview of adaptive parallel reasoning. What if a reasoning model could decide for itself when to decompose and parallelize independent subtasks, how many concurrent threads to spawn, and how to coordinate them based on the problem at hand? We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning. Disclosure: this post is part landscape survey, part perspective on adaptive parallel reasoning. One of the authors (Tony Lian) co-led ThreadWeaver ( Lian et al., 2025 ), one of the methods discussed below. The authors aim to present each approach on its own terms. Motivation Recent progress in LLM reasoning capabilities has been largely driven by inference-time scaling, in addition to data and parameter scaling ( OpenAI et al., 2024 ; DeepSeek-AI et al., 2025 ). Models that explicitly output reasoning tokens (through intermediate steps, backtracking, and exploration) now dominate math, coding, and agentic benchmarks. These behaviors allow models to explore alternative hypotheses, correct earlier mistakes, and synthesize conclusions rather than committing to a single solution ( Wen et al., 2025 ). The problem is that seq

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:Berkeley AI Research · bair.berkeley.edu