跳到正文
Berkeley AI Research·· 2026-07-26

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

中文摘要

ABBEL 与传统递归摘要的对比概述。信念取代了完整的交互历史,作为智能体的工作上下文,信念评分通过监督每个信念状态的内容来提升性能。随着任务时间范围的增长,大型语言模型的上下文无法无限扩展。自我摘要可以生成简洁、可解释的上下文,但会带来显著的性能损失,尤其是在高质量数据稀缺的人类辅助领域,例如协作代码生成。我们通过 ABBEL 解决这一问题:一种框架,以自然语言信念状态的形式隔离并监督摘要的信息内容。动机:递归摘要的成本。为了有效协助日益复杂的任务,如软件开发,语言模型必须能够与我们进行数百甚至数千步的交互。对于如此长的任务,将整个交互历史保留在上下文中是不现实的。迄今为止使用的启发式方法是生成摘要,有时称为上下文压缩。例如,Cursor 最新的模型 Composer 2.5 在训练过程中使用压缩以提高性能(Cassan

英文原文

Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.. As task horizons grow, LLM contexts can’t scale forever. Self-summarization enables concise, interpretable contexts, but at a significant performance cost, especially for human assistance domains where high quality data is scarce, e.g., collaborative code generation. We address this with ABBEL : a framework that isolates and supervises the information content of summaries in the form of natural-language belief states. Motivation: the cost of recursive summarization For language models to effectively assist with increasingly complex tasks such as software development, they must be able to interact with us over hundreds or even thousands of steps. For such long tasks, it is impractical to keep the history of the entire interaction in context. The heuristic approach used so far has been summary generation, sometimes called context compaction. For example, Cursor’s latest model composer 2.5 uses compaction during training for improved performance ( Cassan

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:Berkeley AI Research · bair.berkeley.edu