GlucoFM: Foundation model for continuous glucose monitoring
健康与生物科技
Health & Bioscience
健康与生物科技
Health & Bioscience
人机交互与可视化
Human-Computer Interaction and Visualization
生成式人工智能
Generative AI
算法与理论
Algorithms & Theory
Skala 1.1 是微软研究院推出的更新版深度学习交换相关泛函,它提供了更高的准确性,扩大了在计算化学生态系统中的可访问性,并建立了一个动态的基准来跟踪计算性能。帖子《扩大对 Skala 的访问,为预测 DFT 开辟更快的路径》首先发表在微软研究院。
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research .
通用科学
General Science
一条路径、一堵篱笆、一个结。MindTopo为测试人工智能如何理解拓扑关系设定了新标准,并突显了加强空间推理和规划的新机遇。帖子《MindTopo揭示了VLM的空间推理能力》首先发布在微软研究院。
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
生成式人工智能
Generative AI
健康与生物科学
Health & Bioscience
放射科人工智能正在超越报告生成。CARE-X 探索一种统一的方法,结合灵活的推理、校准的预测和基于测量的工具,用于胸部X光片的解读。帖子《介绍CARE-X:通过辅助监督、奖励对齐学习和工具增强测量,迈向临床有用的放射科视觉语言模型》首先发布在微软研究院。
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .
Orchard 是一个开源框架,供研究界在不同任务类型上训练和评估人工智能代理。通过使研究人员能够重复使用相同的基础设施,Orchard 在减少复杂性的同时,仍能从小型模型中实现强大的性能。文章《Orchard:一个可扩展代理人工智能的开源框架》首先发布于微软研究院。
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research .
通用科学
General Science
图1:CUDA到MLX优化转换图。CUDA优化知识可以转化为原生架构的MLX策略,而不是逐条指令复制。我们正进入计算的新纪元。硬件正在迅速变化——不仅仅是更快的GPU,还有来自不同供应商的越来越多的芯片,每种芯片都有自己的架构,通常针对特定的AI工作负载进行优化。软件也在以同样快的速度变化,如今AI编码工具在几分钟内生成的内容,几年前需要数月的努力才能完成。随着计算的重心如今集中在AI上,GPU内核是其成功的关键组成部分。这些是在GPU内部运行的底层程序,编写高效的内核远非显而易见——需要数年的专业知识才能掌握。将内核从一个供应商的硬件转移到另一个供应商的硬件更加困难,通常意味着要从头开始重新发现相同的优化方法。例如,CUDA生态系统已经积累了数十年来通过艰苦努力获得的内核专业知识:针对注意力机制、状态空间模型和其他关键操作的手动调优实现,这些代表了数千小时的工程努力。较新的硬件生态系统(Apple Silicon,定制AI加速器)
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing. Hardware is changing rapidly — not just faster GPUs, but a growing range of chips from different vendors, each with its own architecture and often tailored to specific AI workloads. Software is changing just as fast, and AI coding tools now generate in minutes what took months of effort a few years ago. With so much of computing now centered on AI, GPU kernels are a crucial component of its success. These are the low-level programs that run inside the GPU, and writing efficient ones is far from obvious — it takes years of expertise to get right. Transferring a kernel from one vendor’s hardware to another is harder still, and often means rediscovering the same optimizations from scratch. The CUDA ecosystem, for example, has accumulated decades of hard-won kernel expertise: hand-tuned implementations of attention, state space models, and other critical operations representing thousands of engineering hours. Newer hardware ecosystems (Apple Silicon, custom AI accelerat
ABBEL 与传统递归摘要的对比概述。信念取代了完整的交互历史,作为智能体的工作上下文,信念评分通过监督每个信念状态的内容来提升性能。随着任务时间范围的增长,大型语言模型的上下文无法无限扩展。自我摘要可以生成简洁、可解释的上下文,但会带来显著的性能损失,尤其是在高质量数据稀缺的人类辅助领域,例如协作代码生成。我们通过 ABBEL 解决这一问题:一种框架,以自然语言信念状态的形式隔离并监督摘要的信息内容。动机:递归摘要的成本。为了有效协助日益复杂的任务,如软件开发,语言模型必须能够与我们进行数百甚至数千步的交互。对于如此长的任务,将整个交互历史保留在上下文中是不现实的。迄今为止使用的启发式方法是生成摘要,有时称为上下文压缩。例如,Cursor 最新的模型 Composer 2.5 在训练过程中使用压缩以提高性能(Cassan
Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.. As task horizons grow, LLM contexts can’t scale forever. Self-summarization enables concise, interpretable contexts, but at a significant performance cost, especially for human assistance domains where high quality data is scarce, e.g., collaborative code generation. We address this with ABBEL : a framework that isolates and supervises the information content of summaries in the form of natural-language belief states. Motivation: the cost of recursive summarization For language models to effectively assist with increasingly complex tasks such as software development, they must be able to interact with us over hundreds or even thousands of steps. For such long tasks, it is impractical to keep the history of the entire interaction in context. The heuristic approach used so far has been summary generation, sometimes called context compaction. For example, Cursor’s latest model composer 2.5 uses compaction during training for improved performance ( Cassan
……民有、民治、民享的政府…… ——亚伯拉罕·林肯,《葛底斯堡演说》(1863年) 人工智能的成本正在迅速下降。2023年初,具备GPT-4级别能力的模型每百万个标记的成本约为30美元;而如今,同样的模型成本已低于1美元,一些供应商甚至将成本压低至0.10美元以下。在各类基准测试中,推理价格每年下降了9倍到900倍不等,中位数下降接近50倍。即使是前沿模型,每一代的成本也在大幅下降,开源模型紧随其后。而且至关重要的是,即使“获得诺贝尔奖的天才级别”的智能尚未出现,目前足以满足绝大多数知识工作的智能已经存在,并且每月都在变得更便宜。以这样的速度,我们即将进入几乎免费的智能时代——这种智能足以满足日常的知识工作。 免责声明:本文由加州大学伯克利分校电子工程与计算机科学系(EECS)副教授、EPIC数据实验室联合主任阿迪蒂亚·G·帕拉梅斯瓦兰(Aditya G. Parameswaran)主导撰写,同时还有他的合作者共同参与。本文部分内容是对现状的综述,部分内容是观点阐述,以下讨论的若干研究方向(包括代理推测)……
... government of the people, by the people, for the people ... — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1 , and some providers are pushing costs below $0.10 . Across benchmarks, inference prices have fallen between 9x and 900x per year , with a median decline near 50x. Even frontier models are getting dramatically cheaper each generation, with open-source models following closely behind. And crucially, even if “Nobel-Prize-winning genius-level” intelligence isn’t here yet, the intelligence that suffices for the vast majority of knowledge work is here today, and getting cheaper by the month. At this rate, we are soon entering the era of virtually free intelligence —the kind that is more than enough for everyday knowledge work. Disclosure: This post is a perspective led by Aditya G. Parameswaran —an Associate Professor of EECS and co-director of the EPIC Data Lab at UC Berkeley—together with his collaborators. It is part landscape survey and part perspective, and several of the research directions discussed below (including agentic speculatio
祝贺伯克利人工智能研究(BAIR)实验室2026届!今年,BAIR迎来了另一组杰出的博士毕业生,他们的好奇心、创造力和毅力推动了人工智能和机器学习的前沿发展。他们的研究涵盖了现代人工智能的广度——从机器人和具身智能、大型语言模型和推理、计算机视觉、生成建模、人工智能安全、人机交互、人工智能在科学和医疗保健中的应用,以及更多领域。在这一过程中,他们发表了有影响力的研究,构建了具有现实世界影响的系统,指导了同龄人,并使BAIR社区变得更好。现在,他们将前往思想传播的任何地方:担任教职和博士后职位、进入工业界研究实验室,或自己创办初创公司——还有一些人仍在探索接下来的路,他们很乐意与您联系。请与我们一同庆祝这些优秀毕业生的成就。我们为他们在伯克利所取得的成就感到自豪,我们迫不及待地想看看他们接下来会做什么!感谢斯坦福人工智能实验室的朋友们提出了这个想法! 白峰石 邮箱:[email protected]
Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning. Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more. Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for the better. Now they are headed everywhere ideas travel: to faculty and postdoctoral positions, to industry research labs, and to startups of their own founding — and several are still exploring what comes next and would love to hear from you. Please join us in celebrating the achievements of these wonderful graduates. We are proud of everything they have accomplished at Berkeley, and we can’t wait to see what they do next! Thank you to our friends at the Stanford AI Lab for this idea! Baifeng Shi Email: [email protected]
自适应并行推理概述。如果一个推理模型可以自行决定何时分解和并行化独立的子任务,生成多少个并发线程,以及如何根据具体问题协调这些线程,会怎样呢?我们对并行推理领域的最新进展进行了详细分析,特别是自适应并行推理。免责声明:本文一部分是对现状的综述,另一部分是对自适应并行推理的见解。其中一位作者(Tony Lian)共同领导了ThreadWeaver(Lian等人,2025年),这是下文讨论的方法之一。作者的目标是根据每种方法自身的特性来呈现它们。动机 近年来,大语言模型推理能力的进展主要由推理时的扩展驱动,除了数据和参数的扩展(OpenAI等,2024;DeepSeek-AI等,2025)。那些显式输出推理标记(通过中间步骤、回溯和探索)的模型现在在数学、编程和代理基准测试中占据主导地位。这些行为使模型能够探索替代假设,纠正早期的错误,并综合得出结论,而不是坚持单一解决方案(Wen等,2025)。问题在于序列
Overview of adaptive parallel reasoning. What if a reasoning model could decide for itself when to decompose and parallelize independent subtasks, how many concurrent threads to spawn, and how to coordinate them based on the problem at hand? We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning. Disclosure: this post is part landscape survey, part perspective on adaptive parallel reasoning. One of the authors (Tony Lian) co-led ThreadWeaver ( Lian et al., 2025 ), one of the methods discussed below. The authors aim to present each approach on its own terms. Motivation Recent progress in LLM reasoning capabilities has been largely driven by inference-time scaling, in addition to data and parameter scaling ( OpenAI et al., 2024 ; DeepSeek-AI et al., 2025 ). Models that explicitly output reasoning tokens (through intermediate steps, backtracking, and exploration) now dominate math, coding, and agentic benchmarks. These behaviors allow models to explore alternative hypotheses, correct earlier mistakes, and synthesize conclusions rather than committing to a single solution ( Wen et al., 2025 ). The problem is that seq
GRASP 是一种基于梯度的新型规划器,用于学习到的动力学(“世界模型”),它通过以下方式使长期规划变得实用:(1) 将轨迹提升到虚拟状态,从而使优化在时间上并行进行;(2) 直接向状态迭代中添加随机性以进行探索;(3) 重新塑造梯度,使动作获得清晰的信号,同时通过高维视觉模型避免脆弱的“状态-输入”梯度。大型的学习世界模型正变得越来越强大。它们可以在高维视觉空间中预测长时间序列的未来观察结果,并以几年前难以想象的方式在不同任务之间进行泛化。随着这些模型的扩展,它们开始看起来不像特定任务的预测器,而更像通用的模拟器。但拥有一个强大的预测模型并不等于能够有效地将其用于控制/学习/规划。实际上,使用现代世界模型进行长期规划仍然很脆弱:优化变得条件不良,非贪心结构会产生不良的局部最小值,而高维潜在空间则引入了微妙的故障模式。在这篇博客文章中,我将描述促使这项研究的问题。
GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models. Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators. But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes. In this blog post, I describe the problems that motivated this pro
理解复杂机器学习系统的行为,尤其是大型语言模型(LLMs),是现代人工智能中的一个关键挑战。可解释性研究旨在使决策过程对模型开发者和受影响的人类更加透明,这是迈向更安全、更可信赖的人工智能的重要一步。为了全面理解这些系统,我们可以从不同的角度进行分析:特征归因(feature attribution),它隔离出推动预测的具体输入特征(Lundberg & Lee, 2017;Ribeiro 等人,2022);数据归因(data attribution),它将模型行为与有影响力的训练样本联系起来(Koh & Liang, 2017;Ilyas 等人,2022);以及机制可解释性(mechanistic interpretability),它剖析内部组件的功能(Conmy 等人,2023;Sharkey 等人,2025)。在这些视角中,同样的基本障碍依然存在:规模上的复杂性。模型行为很少是孤立组件的结果;相反,它源于复杂的依赖关系和模式。为了实现最先进的性能,模型会综合复杂的特征关系,从多样化的训练样本中寻找共同模式,并通过处理信息来实现。
Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decision-making process more transparent to model builders and impacted humans, a step toward safer and more trustworthy AI. To gain a comprehensive understanding, we can analyze these systems through different lenses: feature attribution , which isolates the specific input features driving a prediction ( Lundberg & Lee, 2017 ; Ribeiro et al., 2022 ); data attribution , which links model behaviors to influential training examples ( Koh & Liang, 2017 ; Ilyas et al., 2022 ); and mechanistic interpretability , which dissects the functions of internal components ( Conmy et al., 2023 ; Sharkey et al., 2025 ). Across these perspectives, the same fundamental hurdle persists: complexity at scale . Model behavior is rarely the result of isolated components; rather, it emerges from complex dependencies and patterns. To achieve state-of-the-art performance, models synthesize complex feature relationships, find shared patterns from diverse training examples, and process information throug
编码器(光学系统)将物体映射到无噪声的图像,而噪声会将这些图像转化为测量值。我们的信息估计器仅使用这些带有噪声的测量值和噪声模型,来量化测量值区分物体的能力。许多成像系统产生的测量值人类根本无法看到,或者无法直接理解。你的智能手机在生成最终照片之前,会通过算法处理原始传感器数据。磁共振成像(MRI)扫描仪收集的是频率空间中的测量值,这些测量值需要经过重建后,医生才能查看。自动驾驶汽车会直接通过神经网络处理摄像头和激光雷达(LiDAR)数据。在这些系统中,重要的不是测量值看起来是什么样子,而是它们包含多少有用的信息。即使信息以人类无法理解的方式编码,人工智能也能提取这些信息。然而,我们很少直接评估信息内容。传统的指标如分辨率和信噪比分别评估质量的各个单独方面,这使得在这些因素之间存在权衡的系统之间进行比较变得困难。常见的替代方法是训练神经网络来重建或分类图像,这将成像硬件的质量与神经网络的质量混为一谈。
An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements distinguish objects. Many imaging systems produce measurements that humans never see or cannot interpret directly. Your smartphone processes raw sensor data through algorithms before producing the final photo. MRI scanners collect frequency-space measurements that require reconstruction before doctors can view them. Self-driving cars process camera and LiDAR data directly with neural networks. What matters in these systems is not how measurements look, but how much useful information they contain. AI can extract this information even when it is encoded in ways that humans cannot interpret. And yet we rarely evaluate information content directly. Traditional metrics like resolution and signal-to-noise ratio assess individual aspects of quality separately, making it difficult to compare systems that trade off between these factors. The common alternative, training neural networks to reconstruct or classify images, conflates the quality of the imaging hardware with the qual