Identifying Interactions at Scale for LLMs
Identifying Interactions at Scale for LLMs
理解复杂机器学习系统的行为,尤其是大型语言模型(LLMs),是现代人工智能中的一个关键挑战。可解释性研究旨在使决策过程对模型开发者和受影响的人类更加透明,这是迈向更安全、更可信赖的人工智能的重要一步。为了全面理解这些系统,我们可以从不同的角度进行分析:特征归因(feature attribution),它隔离出推动预测的具体输入特征(Lundberg & Lee, 2017;Ribeiro 等人,2022);数据归因(data attribution),它将模型行为与有影响力的训练样本联系起来(Koh & Liang, 2017;Ilyas 等人,2022);以及机制可解释性(mechanistic interpretability),它剖析内部组件的功能(Conmy 等人,2023;Sharkey 等人,2025)。在这些视角中,同样的基本障碍依然存在:规模上的复杂性。模型行为很少是孤立组件的结果;相反,它源于复杂的依赖关系和模式。为了实现最先进的性能,模型会综合复杂的特征关系,从多样化的训练样本中寻找共同模式,并通过处理信息来实现。
Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decision-making process more transparent to model builders and impacted humans, a step toward safer and more trustworthy AI. To gain a comprehensive understanding, we can analyze these systems through different lenses: feature attribution , which isolates the specific input features driving a prediction ( Lundberg & Lee, 2017 ; Ribeiro et al., 2022 ); data attribution , which links model behaviors to influential training examples ( Koh & Liang, 2017 ; Ilyas et al., 2022 ); and mechanistic interpretability , which dissects the functions of internal components ( Conmy et al., 2023 ; Sharkey et al., 2025 ). Across these perspectives, the same fundamental hurdle persists: complexity at scale . Model behavior is rarely the result of isolated components; rather, it emerges from complex dependencies and patterns. To achieve state-of-the-art performance, models synthesize complex feature relationships, find shared patterns from diverse training examples, and process information throug
应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。
来源:Berkeley AI Research · bair.berkeley.edu