Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
人工智能的新前沿正在发展出能够自主进行研究的科研人员
The new frontier of AI is developing capable autonomous researchers
人工智能的新前沿正在发展出能够自主进行研究的科研人员
The new frontier of AI is developing capable autonomous researchers
通用科学
General Science
一条路径、一堵篱笆、一个结。MindTopo为测试人工智能如何理解拓扑关系设定了新标准,并突显了加强空间推理和规划的新机遇。帖子《MindTopo揭示了VLM的空间推理能力》首先发布在微软研究院。
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
介绍手语转文本(SL2T)模型,这是我们的突破性模型,为聋哑和听力障碍用户的新手语功能提供支持。
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
生成式人工智能
Generative AI
健康与生物科学
Health & Bioscience
放射科人工智能正在超越报告生成。CARE-X 探索一种统一的方法,结合灵活的推理、校准的预测和基于测量的工具,用于胸部X光片的解读。帖子《介绍CARE-X:通过辅助监督、奖励对齐学习和工具增强测量,迈向临床有用的放射科视觉语言模型》首先发布在微软研究院。
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .
Mistral 正在整合推理基础设施、开放模型以及欧洲为掌控其人工智能未来所需做出的长期承诺,并为全球制定了发展路线图。
Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, and setting a roadmap for the world.
你会选择哪个星系?
Which galaxy will you choose?
Shieldstral推出了一种3B参数量的开放权重多模态安全分类器,其性能优于其7倍大小的模型。
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
Orchard 是一个开源框架,供研究界在不同任务类型上训练和评估人工智能代理。通过使研究人员能够重复使用相同的基础设施,Orchard 在减少复杂性的同时,仍能从小型模型中实现强大的性能。文章《Orchard:一个可扩展代理人工智能的开源框架》首先发布于微软研究院。
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research .
我们什么时候建造月球穹顶城市?
When do we build the moon arcology?
通用科学
General Science
Gemini Robotics ER 2 帮助机器人进行推理、协作并解决现实世界中的任务。它标志着在视频理解、工具协调和多机器人协作方面,机器人应用取得了重大进展。
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
图1:CUDA到MLX优化转换图。CUDA优化知识可以转化为原生架构的MLX策略,而不是逐条指令复制。我们正进入计算的新纪元。硬件正在迅速变化——不仅仅是更快的GPU,还有来自不同供应商的越来越多的芯片,每种芯片都有自己的架构,通常针对特定的AI工作负载进行优化。软件也在以同样快的速度变化,如今AI编码工具在几分钟内生成的内容,几年前需要数月的努力才能完成。随着计算的重心如今集中在AI上,GPU内核是其成功的关键组成部分。这些是在GPU内部运行的底层程序,编写高效的内核远非显而易见——需要数年的专业知识才能掌握。将内核从一个供应商的硬件转移到另一个供应商的硬件更加困难,通常意味着要从头开始重新发现相同的优化方法。例如,CUDA生态系统已经积累了数十年来通过艰苦努力获得的内核专业知识:针对注意力机制、状态空间模型和其他关键操作的手动调优实现,这些代表了数千小时的工程努力。较新的硬件生态系统(Apple Silicon,定制AI加速器)
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing. Hardware is changing rapidly — not just faster GPUs, but a growing range of chips from different vendors, each with its own architecture and often tailored to specific AI workloads. Software is changing just as fast, and AI coding tools now generate in minutes what took months of effort a few years ago. With so much of computing now centered on AI, GPU kernels are a crucial component of its success. These are the low-level programs that run inside the GPU, and writing efficient ones is far from obvious — it takes years of expertise to get right. Transferring a kernel from one vendor’s hardware to another is harder still, and often means rediscovering the same optimizations from scratch. The CUDA ecosystem, for example, has accumulated decades of hard-won kernel expertise: hand-tuned implementations of attention, state space models, and other critical operations representing thousands of engineering hours. Newer hardware ecosystems (Apple Silicon, custom AI accelerat
警告枪声将持续到文明醒悟为止
The warning shots will continue until civilization wakes up
ABBEL 与传统递归摘要的对比概述。信念取代了完整的交互历史,作为智能体的工作上下文,信念评分通过监督每个信念状态的内容来提升性能。随着任务时间范围的增长,大型语言模型的上下文无法无限扩展。自我摘要可以生成简洁、可解释的上下文,但会带来显著的性能损失,尤其是在高质量数据稀缺的人类辅助领域,例如协作代码生成。我们通过 ABBEL 解决这一问题:一种框架,以自然语言信念状态的形式隔离并监督摘要的信息内容。动机:递归摘要的成本。为了有效协助日益复杂的任务,如软件开发,语言模型必须能够与我们进行数百甚至数千步的交互。对于如此长的任务,将整个交互历史保留在上下文中是不现实的。迄今为止使用的启发式方法是生成摘要,有时称为上下文压缩。例如,Cursor 最新的模型 Composer 2.5 在训练过程中使用压缩以提高性能(Cassan
Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.. As task horizons grow, LLM contexts can’t scale forever. Self-summarization enables concise, interpretable contexts, but at a significant performance cost, especially for human assistance domains where high quality data is scarce, e.g., collaborative code generation. We address this with ABBEL : a framework that isolates and supervises the information content of summaries in the form of natural-language belief states. Motivation: the cost of recursive summarization For language models to effectively assist with increasingly complex tasks such as software development, they must be able to interact with us over hundreds or even thousands of steps. For such long tasks, it is impractical to keep the history of the entire interaction in context. The heuristic approach used so far has been summary generation, sometimes called context compaction. For example, Cursor’s latest model composer 2.5 uses compaction during training for improved performance ( Cassan