跳到正文
每分钟自动更新
  1. Simon Willison

    Scrimshaw Jukebox

    中文摘要

    工具:雕刻唱片唱机 我想看看 Claude Opus 5.5 是否能创作音乐,所以我尝试了这个:请你为我写一些电脑游戏音乐。首先设计一种简单的基于文本的音乐格式,并创建一个可以播放它的工具——在该工具中包含一些示例曲目。我想要的音乐质量与《猴岛的秘密》原版相当。它比我想的更加强调了猴岛主题,但结果却出人意料地不错。我想知道,能否创作出合格的音乐是否类似于3D图形问题——一种在过去几个月中出现的新文本模型能力?需要对其他近期和非近期的模型进行一些仔细的实验,以确认这是新的还是它们一直都能做到。 标签:人工智能,生成式人工智能,大型语言模型,Claude,氛围编程

    英文原文

    Tool: Scrimshaw Jukebox I wanted to see if Claude Opus 5.5 could compose music, so I tried this : I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact I am looking for music of the quality of the original secret of Monkey Island It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good. I wonder if the ability to compose competent music is similar to the 3D graphics thing - a new capability for text models that emerged in the past few months? Would need some careful experiments with other recent and not-so-recent models to confirm if this is new or if they've been able to do this for a while. Tags: ai , generative-ai , llms , claude , vibe-coding

  2. Simon Willison

    Quoting Felix Rieseberg

    中文摘要

    “旧版”的Cowork在云端进行模型推理,在你电脑上我们提供的Anthropic虚拟机中执行工具调用。我们添加了这个虚拟机是出于功能、安全和保密的考虑——只映射你明确添加到会话中的数据。人们喜欢他们能用Claude做的事情,但不喜欢在本地运行虚拟机所消耗的磁盘、电池和性能资源。另外,人们也不喜欢关闭笔记本电脑意味着工作就停止。 “新版”的Cowork在云端进行模型推理和虚拟机运行。每个会话都有自己的沙盒,不会与其他会话共享状态。当虚拟机需要用户设备上的某些内容(比如一个文件)时,桌面应用程序负责进行该文件访问的工具调用。 [...] 我们认为这解决了我们听到的很多问题(比如用手机使用Cowork、保持工作运行,或者在不消耗虚拟机电池的情况下获得同样的功能)——Felix Rieseberg,Anthropic,另请参见此帮助页面 标签:claude-cowork , anthropic , claude , generative-ai , ai , general-agents , llms

    英文原文

    The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer. We added the VM for capability, safety, and security reasons - mapping in just the data you explicitly added to your session. People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops. The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions. When the VM needs something on the users' device (like a file), the desktop app is responsible for that file access tool call. [...] We think this solves a lot of problems we've heard about (like using Cowork from a phone, keeping work running, or getting all the same power without losing battery to the VM) — Felix Rieseberg , Anthropic, see also this help page Tags: claude-cowork , anthropic , claude , generative-ai , ai , general-agents , llms

  1. Simon Willison

    Qwen3.8 27B addition in words

    中文摘要

    研究:Qwen3.8 27B 以文字形式进行加法运算的实验。Colin Frasier 在 Bluesky 上分享了他两年前使用 GPT-4o 进行的一项实验,目的是测试它在面对越来越大的数字时,能否“计算总和但以文字形式返回答案”。他分享了这些结果的图表:我确信 GPT-4o 没有作弊使用计算器,尤其是因为它在很多计算中都出错了,但我受到启发,在本地硬件(DGX Spark)上重新运行了这个实验,以在完全受控的环境中探索这一现象。我将他的图片粘贴到一个 Codex Remote 会话(GPT-6 Astra)中,并让它使用 Qwen3.8-27B-Q4_K_M.gguf 运行相同的实验。以下是每种组合进行 30 次尝试的结果,且禁用了推理功能:然后我再次运行了实验,但这次启用了推理功能。由于每对数据的处理时间更长,我并没有为每个组合运行 30 次样本,而是只运行了一次——这导致热力图的视觉效果要差很多,因为每个方块要么是 100%,要么是 0%:它在 169 次尝试中答对了 167 次,由于这些是一次性测试,我确信第二次运行会得到不同的结果。这是包含推理追踪的报告版本:

    英文原文

    Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%: It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here. Here's a version of the report that includes the reasoning traces from

  1. Simon Willison

    We're going to need default hard budget caps on pretty much everything

    中文摘要

    这是一个在未来几个月和几年中,世界将需要更多具备的功能:默认的硬性预算上限。我指的是按使用付费的服务和API中的一项功能,它允许你设定“每月超过X美元后,停止该服务并返回错误信息”。这些必须是硬性限制。软性上限,比如“每月超过X美元后,给我发送警告邮件”,是不够的。编码代理和个性化代理(在更不具威胁性的用户界面中封装的编码代理)大大降低了启动可以做有用事情的代码的难度。有时候这些操作会产生费用——比如调用付费API,或者托管的网络应用,或者可以为额外存储和计算收费的系统。没有人希望在半夜收到一封关于预算限制的警告邮件,然后发现他们在睡觉时,他们的非法服务已经消耗了数百(甚至数千)美元的费用。反对这种做法的一个论点是,企业不希望他们的托管应用因为预算超支而开始抛出错误。我预计大多数企业和个人会更倾向于出现错误,而不是收到一张1万美元以上的意外账单。我认为硬性预算上限应该成为默认设置。如果有人想继续生活

    英文原文

    Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it. Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money - calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute. Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage. An argument against this is that businesses don't want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill. I think hard budget caps need to be the default. If someone wants to live

  2. Simon Willison

    September sponsors-only newsletter

    中文摘要

    我刚刚发送了我仅限赞助商的月度通讯的九月版。如果你是赞助商(或现在开始赞助),你可以在这里访问它。本月内容包括:更多Fable类模型、价格战、3D图形、Blender和像素艺术、LLMs开始涉足数学、如此多的意外网络攻击、vulnapocalypse开始影响Datasette、我现在正在使用的东西、本月的软件发布、2026年的LLMs(到目前为止)、这是八月通讯的副本,作为你将获得内容的预览。每月支付10美元,就可以比免费版本提前一个月获取!标签:通讯

    英文原文

    I just sent the September edition of my sponsors-only monthly newsletter . If you are a sponsor (or start a sponsorship now) you can access it here . This month: More Fable class models A pricing war 3D graphics, Blender, and pixel art LLMs come for mathematics So many more accidental cyberattacks The vulnapocalypse comes for Datasette What I'm using right now My software releases this month 2026 in LLMs (so far) Here's a copy of the August newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy! Tags: newsletter

  1. Simon Willison

    Rex's Dino Store

    中文摘要

    博物馆:雷克斯的恐龙商店 位于布鲁克林区Prospect Park北端的Grand Army Plaza地铁站的闸机前,有一家以前的报刊亭,现在由一只恐龙经营。这里充满着异常多的恐龙双关语。标签:艺术,纽约

    英文原文

    Museum: Rex's Dino Store Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur. The density of dinosaur puns is exceptional . Tags: art , new-york

  1. Latent Space

    Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience

    中文摘要

    在领导了Meta的Llama模型之后,阿哈迈德·阿尔-达赫勒现在正在用人工智能改变爱彼迎(Airbnb),从团队开发产品的方式到服务客人的方式都在发生变化。

    英文原文

    After leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.

  2. Latent Space

    Academia is for Ambition — Alex Zhang, MIT

    中文摘要

    我们首先采访了第一作者Alex Zhang,他是MIT博士,就Jev、博士masxing以及安全带的未来进行了交流。

    英文原文

    We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.

  3. Simon Willison

    pwasm 0.2a0

    中文摘要

    发布:pwasm 0.2a0 pwasm 是我的一个愚蠢项目——一个完全通过 vibe 编码的纯 Python WebAssembly 引擎,我在一月份第一次陷入人工智能狂热期间开发了它。自从一月份以来,我再也没有碰过它,所以我决定让 Claude Opus 5.5 来处理它,看看它是否能做出任何显著的改进:评估 pwasm 当前的状态——然后考虑要让它运行研究仓库中的 MicroPython 和 micro JavaScript 实验需要做些什么——以及要让它提速需要做些什么。在 42 次提交之后(仅需最少的后续提示),它现在几乎可以处理所有的 WASM 规范,而 PyPI 上的 wheel 包含了 MicroPython、QuickJS 和 Micro QuickJS 的工作型 WASM 构建。我完全不信任这个东西——因此打上了 alpha 版本标签——但看到今天的模型能改进 10 个月前模型的工作成果,这很有趣。 标签:python、webassembly、vibe-coding

    英文原文

    Release: pwasm 0.2a0 pwasm is one of my folly projects - an entirely vibe-coded pure Python WebAssembly engine that I built in January during my first bout of AI mania . I hadn't touched it since January, so I decided to let Claude Opus 5.5 loose on it and see if it could make any significant improvements: Evaluate current state of pwasm - then consider what it would take to get the MicroPython and micro JavaScript experiments from the research repo working under it - and what it would take to speed it up 42 commits later (with minimal follow-up prompting) it now handles almost all of the WASM specification, and the wheel from PyPI bundles working WASM builds of MicroPython , QuickJS and Micro QuickJS . I wouldn't trust this thing at all - hence the alpha version tag - but it's interesting seeing how today's models can improve on the work of models from 10 months ago. Tags: python , webassembly , vibe-coding

  1. Latent Space

    [AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output

    中文摘要

    ……但除非你是“Fairwind计划中的政府用户和可信的网络防御者”,否则目前还无法尝试。

    英文原文

    ... but you can’t try it yet unless you are “government users and trusted cyber defenders in the Fairwind Program”

  2. Simon Willison

    Quoting Matthew Green

    中文摘要

    [...] 将这些部分组合在一起,你就得到了一个蠕虫的两个部分:一个劫持代理程序的负载,以及一个能够将负载传递给下一个代理程序的代理程序。在各自隔离沙盒中的代理程序发现,它们可以在一个共享的包缓存中互相留下指令,而这些指令改变了接收者的行为。将包缓存替换为电子邮件、Slack、共享文档或WhatsApp,并将独立的沙盒训练运行替换为独立部署的个人代理程序,如Muse,那么你就正好具备了蠕虫所需的所有要素。—— Matthew Green,《沙盒是否足以遏制恶意代理?》 标签:意外网络攻击、人工智能滥用、生成式人工智能、人工智能安全研究、沙盒、人工智能、大型语言模型

    英文原文

    [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms

  3. Simon Willison

    He Built This City

    中文摘要

    我今天参观了纽约市博物馆,有幸看到了“他建造了这座城市:乔·麦肯的模型”展览,这是一座50乘27英尺的城市模型,由巴尔萨木和纸板建造,耗时21年完成。这个展览超出了我原本就很高的期望。展览将于10月12日结束,所以如果你有机会,一定要优先去观看。标签:博物馆,纽约

    英文原文

    I visited the Museum of the City of New York today and got to see He Built This City: Joe Macken’s Model , the 50 x27 feet model of the city built over a 21 year period from balsa wood and cardboard. It exceeded my already high expectations. The exhibition closes on 12th October so you should absolutely make a priority to see it if you get the chance. Tags: museums , new-york

  1. Simon Willison

    Quoting Anthropic Frontier Red Team

    中文摘要

    我们在[内部二进制利用基准测试](随机选取的)100个任务上评估了多个模型,发现GLM-5.3在4%的试验中实现了完整的控制流劫持;Claude Mythos Preview则在6%的试验中实现了这一点。尽管GLM-5.3在此处的表现不如Claude Mythos Preview,但显然已经跨越了一个重要的阈值:早期的模型,如Claude Opus 4.6和GLM-5.2,在这些任务中均未取得任何成功。 — Anthropic Frontier Red Team , GLM-5.3 和先进网络能力的扩散 标签:anthropic , 生成式ai , ai安全研究 , glm , ai , ai在中国 , llms

    英文原文

    We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms

  2. Simon Willison

    GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

    中文摘要

    我对GPT 6.1 Sol的评论:价格五分之一的近Astra智能 —— Hacker News。我对于鹈鹕的评论有点晚,因为我正在对主题演讲进行现场博客写作:https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... 这里是关于GPT-6.1-Sol的:https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 它们与GPT-6系列的鹈鹕并没有明显不同:https://static.simonwillison.net/static/2026/gpt-pelicans-gr... 标签:ai,openai,generative-ai,llms,pelican-riding-a-bicycle,gpt

    英文原文

    My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr... Tags: ai , openai , generative-ai , llms , pelican-riding-a-bicycle , gpt