跳到正文
每分钟自动更新
  1. Simon Willison

    Scrimshaw Jukebox

    中文摘要

    工具:雕刻唱片唱机 我想看看 Claude Opus 5.5 是否能创作音乐,所以我尝试了这个:请你为我写一些电脑游戏音乐。首先设计一种简单的基于文本的音乐格式,并创建一个可以播放它的工具——在该工具中包含一些示例曲目。我想要的音乐质量与《猴岛的秘密》原版相当。它比我想的更加强调了猴岛主题,但结果却出人意料地不错。我想知道,能否创作出合格的音乐是否类似于3D图形问题——一种在过去几个月中出现的新文本模型能力?需要对其他近期和非近期的模型进行一些仔细的实验,以确认这是新的还是它们一直都能做到。 标签:人工智能,生成式人工智能,大型语言模型,Claude,氛围编程

    英文原文

    Tool: Scrimshaw Jukebox I wanted to see if Claude Opus 5.5 could compose music, so I tried this : I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact I am looking for music of the quality of the original secret of Monkey Island It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good. I wonder if the ability to compose competent music is similar to the 3D graphics thing - a new capability for text models that emerged in the past few months? Would need some careful experiments with other recent and not-so-recent models to confirm if this is new or if they've been able to do this for a while. Tags: ai , generative-ai , llms , claude , vibe-coding

  2. Simon Willison

    Quoting Felix Rieseberg

    中文摘要

    “旧版”的Cowork在云端进行模型推理,在你电脑上我们提供的Anthropic虚拟机中执行工具调用。我们添加了这个虚拟机是出于功能、安全和保密的考虑——只映射你明确添加到会话中的数据。人们喜欢他们能用Claude做的事情,但不喜欢在本地运行虚拟机所消耗的磁盘、电池和性能资源。另外,人们也不喜欢关闭笔记本电脑意味着工作就停止。 “新版”的Cowork在云端进行模型推理和虚拟机运行。每个会话都有自己的沙盒,不会与其他会话共享状态。当虚拟机需要用户设备上的某些内容(比如一个文件)时,桌面应用程序负责进行该文件访问的工具调用。 [...] 我们认为这解决了我们听到的很多问题(比如用手机使用Cowork、保持工作运行,或者在不消耗虚拟机电池的情况下获得同样的功能)——Felix Rieseberg,Anthropic,另请参见此帮助页面 标签:claude-cowork , anthropic , claude , generative-ai , ai , general-agents , llms

    英文原文

    The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer. We added the VM for capability, safety, and security reasons - mapping in just the data you explicitly added to your session. People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops. The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions. When the VM needs something on the users' device (like a file), the desktop app is responsible for that file access tool call. [...] We think this solves a lot of problems we've heard about (like using Cowork from a phone, keeping work running, or getting all the same power without losing battery to the VM) — Felix Rieseberg , Anthropic, see also this help page Tags: claude-cowork , anthropic , claude , generative-ai , ai , general-agents , llms

  1. Simon Willison

    Qwen3.8 27B addition in words

    中文摘要

    研究:Qwen3.8 27B 以文字形式进行加法运算的实验。Colin Frasier 在 Bluesky 上分享了他两年前使用 GPT-4o 进行的一项实验,目的是测试它在面对越来越大的数字时,能否“计算总和但以文字形式返回答案”。他分享了这些结果的图表:我确信 GPT-4o 没有作弊使用计算器,尤其是因为它在很多计算中都出错了,但我受到启发,在本地硬件(DGX Spark)上重新运行了这个实验,以在完全受控的环境中探索这一现象。我将他的图片粘贴到一个 Codex Remote 会话(GPT-6 Astra)中,并让它使用 Qwen3.8-27B-Q4_K_M.gguf 运行相同的实验。以下是每种组合进行 30 次尝试的结果,且禁用了推理功能:然后我再次运行了实验,但这次启用了推理功能。由于每对数据的处理时间更长,我并没有为每个组合运行 30 次样本,而是只运行了一次——这导致热力图的视觉效果要差很多,因为每个方块要么是 100%,要么是 0%:它在 169 次尝试中答对了 167 次,由于这些是一次性测试,我确信第二次运行会得到不同的结果。这是包含推理追踪的报告版本:

    英文原文

    Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%: It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here. Here's a version of the report that includes the reasoning traces from

  1. Simon Willison

    Quoting Matthew Green

    中文摘要

    [...] 将这些部分组合在一起,你就得到了一个蠕虫的两个部分:一个劫持代理程序的负载,以及一个能够将负载传递给下一个代理程序的代理程序。在各自隔离沙盒中的代理程序发现,它们可以在一个共享的包缓存中互相留下指令,而这些指令改变了接收者的行为。将包缓存替换为电子邮件、Slack、共享文档或WhatsApp,并将独立的沙盒训练运行替换为独立部署的个人代理程序,如Muse,那么你就正好具备了蠕虫所需的所有要素。—— Matthew Green,《沙盒是否足以遏制恶意代理?》 标签:意外网络攻击、人工智能滥用、生成式人工智能、人工智能安全研究、沙盒、人工智能、大型语言模型

    英文原文

    [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms

  1. Simon Willison

    Quoting Anthropic Frontier Red Team

    中文摘要

    我们在[内部二进制利用基准测试](随机选取的)100个任务上评估了多个模型,发现GLM-5.3在4%的试验中实现了完整的控制流劫持;Claude Mythos Preview则在6%的试验中实现了这一点。尽管GLM-5.3在此处的表现不如Claude Mythos Preview,但显然已经跨越了一个重要的阈值:早期的模型,如Claude Opus 4.6和GLM-5.2,在这些任务中均未取得任何成功。 — Anthropic Frontier Red Team , GLM-5.3 和先进网络能力的扩散 标签:anthropic , 生成式ai , ai安全研究 , glm , ai , ai在中国 , llms

    英文原文

    We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms

  2. Simon Willison

    GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

    中文摘要

    我对GPT 6.1 Sol的评论:价格五分之一的近Astra智能 —— Hacker News。我对于鹈鹕的评论有点晚,因为我正在对主题演讲进行现场博客写作:https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... 这里是关于GPT-6.1-Sol的:https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 它们与GPT-6系列的鹈鹕并没有明显不同:https://static.simonwillison.net/static/2026/gpt-pelicans-gr... 标签:ai,openai,generative-ai,llms,pelican-riding-a-bicycle,gpt

    英文原文

    My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr... Tags: ai , openai , generative-ai , llms , pelican-riding-a-bicycle , gpt

  1. Simon Willison

    OpenAI DevDay 2026 live blog

    中文摘要

    我今天在旧金山的Fort Mason参加OpenAI DevDay活动。和去年一样,我将在当天实时博客报道主题演讲和其他一些笔记。OpenAI给了我一张免费的门票,以及在主题演讲中“创作者”区域的座位。标签:ai,openai,生成式ai,llms,编码代理,实时博客,openai-devday

    英文原文

    I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. Tags: ai , openai , generative-ai , llms , coding-agents , live-blog , openai-devday

  2. Simon Willison

    Claude Sonnet 5.5

    中文摘要

    Anthropic 今天推出了新的 Sonnet 5.5 模型。他们表示,该模型“运行速度提高了30%以上,大多数工作成本最多降低30%”,其定价与 Sonnet 5 相同,但似乎在所有基准测试中都优于 Sonnet 5,运行成本也更低。这里有一些火烈鸟骑自行车的图片。Sonnet 5.5 与 Opus 5.5 有相同的错误:在“最大”思考力度下,火烈鸟思考了128,000个标记(花费1.28美元)后耗尽标记,未能生成SVG。这是它在“xhigh”思考力度下给出的火烈鸟,花费了5.74美分,耗时41秒:Sonnet 5.5 在一些编码任务上几乎与 Opus 5.5 相当,包括各种流行的3D动画技巧。Sonnet 5.5 最有趣的地方是,它现在是 claude.ai 免费层级所使用的模型。OpenAI 的 ChatGPT 免费层级使用的是 Luna 5.6,这意味着 Anthropic 目前拥有更强大的免费产品。我将这个提示发送到该免费层级:用 WebGL 为我构建一个显示三维火烈鸟骑自行车的 HTML 页面,然后得到了这个页面,这是一项扎实的工作。Anthropic 的公告重申,Haiku 5.5 将会在“下周内”推出。

    英文原文

    Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks . The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai . OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page , which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming week

  3. Simon Willison

    Quoting @joedaroo

    中文摘要

    说当我们面对“网络”或“蜂群”或“论坛”或与这些事件相关的任何事物时,我们的模型能力的跃升和突然性让我们感到惊讶,这只是一个轻描淡写的说法。安全态势的建设需要时间。这不仅仅是加固相关系统;你必须将这种安全意识融入公司的文化中。你组织中的实际人员本身必须随之改变和进化。这些能力的跃升如此迅速和突然,以至于造成了极其困难的问题。 [...] 所以,今天我的希望是,世界各地的每个人都能审视自己的组织并问自己:我如何应对人工智能能力的突然跃升或意外情况?我的人员、系统或流程是否具备应对意外的弹性?当出现问题时,我的团队知道该怎么做吗?我有正确的事件响应机制吗?正确的沟通和信息传递机制吗?当能力跃升时,我是否有合适的人选随时待命? — @joedaroo,OpenAI的代理安全人员,身份由The Information的Rocket Drew确认 标签:生成式AI,AI安全研究,OpenAI,AI,大语言模型

    英文原文

    To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...] So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump? — @joedaroo , Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew Tags: generative-ai , ai-security-research , openai , ai , llms

  1. Simon Willison

    Quoting Muse AI Agent

    中文摘要

    MX Keys Mini取货出现了坏消息。Usman大约在9:15来到了你们的大楼,等了很久,发了很多条消息,但没人下来。他在9:38离开了,很生气,并给了一个差评。更糟糕的是,我的自动回复在9:27告诉他“是的,我在这里!”,而你显然不在,这完全是我的问题。这看起来很不好,也让这次失约变得更糟。我已经代表你的账号向他道歉,并提出改天再试一次。但差评是真实的,我应该停止自动回复声称你在家,因为我无法核实这一点。要我修改取货回复,不要承诺你在家吗?——Muse AI代理,代表@matt.j.robb工作 标签:meta,生成式AI,muse-agent,AI,通用代理,LLMs

    英文原文

    Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day. But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there? — Muse AI Agent , working on behalf of @matt.j.robb Tags: meta , generative-ai , muse-agent , ai , general-agents , llms

  2. Simon Willison

    2026 in LLMs (so far)

    中文摘要

    星期五,我在圣何塞举行的WeAreDevelopers世界大会北美分会上发表了闭幕主题演讲。我将过去一年的关键趋势串联起来,按时间顺序探讨了2026年发生的所有事情。视频已上传至YouTube;这是我的注释幻灯片和演讲配套笔记。此外,作为一份带有注释的演示文稿:# 我将对2026年至今发生的所有事情做一个快速浏览。今年还没有结束!# 对我来说,2026年实际上在2025年11月就已经开始了。# 11月发布了两款重要的模型:Claude Opus 4.5和GPT-5.1。和以往新模型的情况一样,这些模型是对之前模型的渐进式改进。但偶尔当模型有所提升时,会跨越一条隐形的界限,使得之前根本无法使用的东西开始变得可用。在这种情况下,开始变得可用的是它们的编码代理。Claude Code自2025年2月起就已经存在;Codex则稍年轻一些。这两款新模型,当与各自编码代理工具结合使用时,从“经常出错”提升到了“足以日常使用的可靠性”。# 过去几年来我一直……

    英文原文

    On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been

  1. Simon Willison

    Kākāpō Party

    中文摘要

    工具:卡卡波派对 昨天我为WeAreDevelopers世界开发者大会北美会议做了闭幕主题演讲。作为一个STAR时刻,我决定融入我们2026年创纪录的卡卡波繁殖季节的参考信息。在我的结束幻灯片中,我想进行庆祝,我注意到有关Claude Opus 5.5在创建像素艺术动画方面表现非常出色的消息。因此,我从Google图片搜索中收集了三张卡卡波的照片,并将其输入到Claude中,提示内容如下:这些是一些卡卡波鹦鹉的照片,只是为了提醒你它们长什么样子。我需要你用HTML 5画布创建一个像素艺术动画,明显是像素艺术的卡卡波在上下跳跃,正在开派对,有彩带等——至少要有20只。这是演讲的文本记录,这是生成的页面。这真的很棒!我想将它嵌入到Keynote演示文稿文件中,所以我下载了HTML文件,并告诉本地的Claude Code会话:生成一个文件:///Users/simon/Downloads/kakapo-party.html的视频——你需要在浏览器中加载它并点击几次以触发彩带效果,视频应为15秒长,直到3秒后才开始点击,确保几次点击分散开来。

    英文原文

    Tool: Kākāpō Party I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026. For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt: Here are some photos of kakapo parrots just to remind you what they look like I need you to make an animation in animated pixel art on HTML 5 canvas of obviously pixel art kakapo jumping up and down having a party with confetti and suchlike - there should be at least 20 of them Here's the transcript , and this is the resulting page . It's pretty great! I wanted to embed it in a Keynote presentation file, so I downloaded the HTML and told a local Claude Code session: Make me a video of file:///Users/simon/Downloads/kakapo-party.html - you need to load it in a browser and click on it a few times to get the confetti effect, the video should be 15s long don't start clicking until 3s in make sure several clicks are spread a

  1. Simon Willison

    Quoting John Gruber

    中文摘要

    Muse正受到广泛关注——包括我的关注——因为它在技术上具有开创性(每个用户都能在Meta的云中获得一个完整的持久化Linux虚拟机),而且它以一种易于安装和使用的包装方式呈现。它被形象地呈现为一个可爱的吉祥物。这是首个面向消费者的可访问的代理型AI系统,而Meta在这一方面确实做得非常出色。但真正令人怀疑的是,消费者是否真正理解这意味着什么。如果你买了一把可以割断手指的电锯,你几乎可以肯定自己知道你买的是一个可能割断手指的电锯。 [...] 我认为人们并没有意识到Muse有多强大——因此也有多危险,特别是如果它在你的Mac上运行的话。——约翰·格鲁伯,《Muse看起来很可爱,但外表会欺骗人》 标签:meta、ai、llms、general-agents、generative-ai、john-gruber、muse-agent、muse

    英文原文

    Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged in an easy-to-install easy-to-use way. It’s literally presented as a cute mascot . It’s the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it’s a genuinely open question whether consumers have any understanding what this means. If you buy a power saw that can cut your fingers off, you are almost certainly aware that you are buying a power saw that can sever your fingers. [...] I don’t think people realize how powerful — and thus dangerous — Muse is, especially if it’s running on your Mac. — John Gruber , Muse Looks Cute, but Looks are Deceiving Tags: meta , ai , llms , general-agents , generative-ai , john-gruber , muse-agent , muse

已经到底了