2026 in LLMs (so far)
星期五,我在圣何塞举行的WeAreDevelopers世界大会北美分会上发表了闭幕主题演讲。我将过去一年的关键趋势串联起来,按时间顺序探讨了2026年发生的所有事情。视频已上传至YouTube;这是我的注释幻灯片和演讲配套笔记。此外,作为一份带有注释的演示文稿:# 我将对2026年至今发生的所有事情做一个快速浏览。今年还没有结束!# 对我来说,2026年实际上在2025年11月就已经开始了。# 11月发布了两款重要的模型:Claude Opus 4.5和GPT-5.1。和以往新模型的情况一样,这些模型是对之前模型的渐进式改进。但偶尔当模型有所提升时,会跨越一条隐形的界限,使得之前根本无法使用的东西开始变得可用。在这种情况下,开始变得可用的是它们的编码代理。Claude Code自2025年2月起就已经存在;Codex则稍年轻一些。这两款新模型,当与各自编码代理工具结合使用时,从“经常出错”提升到了“足以日常使用的可靠性”。# 过去几年来我一直……
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been