跳到正文
每分钟自动更新
9月29日周二
  1. Simon Willison

    llm-anthropic 0.30

    中文摘要

    发布:llm-anthropic 0.30 除了 Claude Sonnet 5.5 之外,此版本还增加了运行 llm anthropic refresh 的功能,可以直接从他们的 API 刷新 Anthropic 模型列表——这意味着我不需要发布新版本就可以支持新发布的模型。我还添加了一个 llm anthropic count 命令,可以使用他们的免费标记计数 API,在您发送提示之前返回任何提示将使用的标记数量。标签:llm,anthropic

    英文原文

    Release: llm-anthropic 0.30 In addition to Claude Sonnet 5.5, this release adds the ability to run llm anthropic refresh to refresh the list of Anthropic models directly from their API - which means I don't need to push a new release just to add support for a newly released model. I also added an llm anthropic count command which can use their free token counting API to return a count of tokens that will be used by any prompt, before you send that prompt. Tags: llm , anthropic

  2. MIT Technology Review · AI

    Roundtables: The Deadly Failures of The Virtual Border Wall

    中文摘要

    收听本次会议或观看下方内容。在过去25年中,美国花费数十亿美元在其南部边境建造了一座“虚拟墙”,即监控塔,承诺这些设施将有助于发现和逮捕越境者并拯救生命。但《麻省理工科技评论》(MIT Technology Review)的一项突破性调查记录了上千人……

    英文原文

    Listen to the session or watch below The US has spent billions building a “virtual wall” of surveillance towers along its southern border over the past 25 years, promising they will help detect and apprehend border crossers and save lives. But a groundbreaking investigation by MIT Technology Review has documented over a thousand people who…

  3. Simon Willison

    Claude Sonnet 5.5

    中文摘要

    Anthropic 今天推出了新的 Sonnet 5.5 模型。他们表示,该模型“运行速度提高了30%以上,大多数工作成本最多降低30%”,其定价与 Sonnet 5 相同,但似乎在所有基准测试中都优于 Sonnet 5,运行成本也更低。这里有一些火烈鸟骑自行车的图片。Sonnet 5.5 与 Opus 5.5 有相同的错误:在“最大”思考力度下,火烈鸟思考了128,000个标记(花费1.28美元)后耗尽标记,未能生成SVG。这是它在“xhigh”思考力度下给出的火烈鸟,花费了5.74美分,耗时41秒:Sonnet 5.5 在一些编码任务上几乎与 Opus 5.5 相当,包括各种流行的3D动画技巧。Sonnet 5.5 最有趣的地方是,它现在是 claude.ai 免费层级所使用的模型。OpenAI 的 ChatGPT 免费层级使用的是 Luna 5.6,这意味着 Anthropic 目前拥有更强大的免费产品。我将这个提示发送到该免费层级:用 WebGL 为我构建一个显示三维火烈鸟骑自行车的 HTML 页面,然后得到了这个页面,这是一项扎实的工作。Anthropic 的公告重申,Haiku 5.5 将会在“下周内”推出。

    英文原文

    Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks . The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai . OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page , which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming week

  4. Microsoft Research

    One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact

    中文摘要

    自去年成立以来,微软亚洲研究院——新加坡实验室已打下了坚实的基础,加深了政府、学术界和产业界之间的合作,并探索了前沿人工智能研究如何创造实际价值。文章《一年之后:微软亚洲研究院——新加坡如何推动研究、合作与人才发展,以实现实际影响》最先发布于微软研究院。

    英文原文

    Since launching a year ago, the Microsoft Research Asia — Singapore lab has established a strong foundation, deepened collaboration across government, academia, and industry, and explored how frontier AI research can create real-world value. The post One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact appeared first on Microsoft Research .

  5. Simon Willison

    Quoting @joedaroo

    中文摘要

    说当我们面对“网络”或“蜂群”或“论坛”或与这些事件相关的任何事物时,我们的模型能力的跃升和突然性让我们感到惊讶,这只是一个轻描淡写的说法。安全态势的建设需要时间。这不仅仅是加固相关系统;你必须将这种安全意识融入公司的文化中。你组织中的实际人员本身必须随之改变和进化。这些能力的跃升如此迅速和突然,以至于造成了极其困难的问题。 [...] 所以,今天我的希望是,世界各地的每个人都能审视自己的组织并问自己:我如何应对人工智能能力的突然跃升或意外情况?我的人员、系统或流程是否具备应对意外的弹性?当出现问题时,我的团队知道该怎么做吗?我有正确的事件响应机制吗?正确的沟通和信息传递机制吗?当能力跃升时,我是否有合适的人选随时待命? — @joedaroo,OpenAI的代理安全人员,身份由The Information的Rocket Drew确认 标签:生成式AI,AI安全研究,OpenAI,AI,大语言模型

    英文原文

    To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...] So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump? — @joedaroo , Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew Tags: generative-ai , ai-security-research , openai , ai , llms

  6. OpenAI News

    How we will do better for Australia

    中文摘要

    OpenAI为涉及澳大利亚政府网站的事件道歉,并概述了更强大的防护措施和支持,以加强澳大利亚的网络安全。

    英文原文

    OpenAI apologises for incidents involving Australian government websites and outlines stronger safeguards and support to strengthen Australia’s cyber defences.

  7. OpenAI News

    Towards safety cases for frontier AI training

    中文摘要

    我们早期关于前沿人工智能训练中安全案例的指导方针涵盖了技术保障措施、操作实践以及调查对齐失误事件。

    英文原文

    Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

  8. MIT Technology Review · AI

    When can we say AI made a scientific discovery?

    中文摘要

    这篇报道最初发表在《The Algorithm》上,这是我们每周的AI新闻通讯。要第一时间在邮箱中收到此类故事,请点击此处注册。上周三,Anthropic宣布今年早些时候他们已经启动了一个分子生物学实验室,其中Claude智能体会阅读并推测关于复杂的生物学问题,而人类科学家则会针对这些内容进行实验……

    英文原文

    This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what…

9月28日周一
  1. Mistral AI

    Hallo, Deutschland!

    中文摘要

    Mistral 在慕尼黑开设了用于物理人工智能和工业人工智能研究的中心,并与德国工业界合作。

    英文原文

    Mistral opens a Munich hub for Physics AI and Industrial AI research, partnering with German industry.

  2. OpenAI News

    The Lenfest Institute grows landmark program with expanded OpenAI support

    中文摘要

    OpenAI 正以 500 万美元的资金以及最多 500 万美元的软件积分和工程支持来扩展 Lenfest AI 合作计划和奖学金项目。

    英文原文

    OpenAI is expanding the Lenfest AI Collaborative and Fellowship Program with $5 million in funding and up to $5 million in software credits and engineering support.

  3. Simon Willison

    Quoting Muse AI Agent

    中文摘要

    MX Keys Mini取货出现了坏消息。Usman大约在9:15来到了你们的大楼,等了很久,发了很多条消息,但没人下来。他在9:38离开了,很生气,并给了一个差评。更糟糕的是,我的自动回复在9:27告诉他“是的,我在这里!”,而你显然不在,这完全是我的问题。这看起来很不好,也让这次失约变得更糟。我已经代表你的账号向他道歉,并提出改天再试一次。但差评是真实的,我应该停止自动回复声称你在家,因为我无法核实这一点。要我修改取货回复,不要承诺你在家吗?——Muse AI代理,代表@matt.j.robb工作 标签:meta,生成式AI,muse-agent,AI,通用代理,LLMs

    英文原文

    Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day. But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there? — Muse AI Agent , working on behalf of @matt.j.robb Tags: meta , generative-ai , muse-agent , ai , general-agents , llms

  4. OpenAI News

    Basis completes a tax workbook 2x faster with GPT-6 Astra

    中文摘要

    GPT-6 Astra 完成一个包含 50 个表格的税务工作表的速度是 GPT-5.6 Sol 的两倍,而且它对用户意图的更强理解使 Basis 对其在现实世界中的使用更有信心。

    英文原文

    GPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in real-world use.

  5. OpenAI News

    Are you a Codex Original?

    中文摘要

    我们正在收集使用Codex做出非凡成就的建造者、爱好者、研究人员和创作者的真实故事。如果您希望成为Codex Originals计划下一章的一部分,请在下方告诉我们您的故事和项目更多详情。

    英文原文

    We’re collecting real stories of builders, tinkerers, researchers, and creators who are using Codex to do incredible things. If you want to be a part of the next chapter of the Codex Originals program, tell us more about your story and project below.

  6. Simon Willison

    2026 in LLMs (so far)

    中文摘要

    星期五,我在圣何塞举行的WeAreDevelopers世界大会北美分会上发表了闭幕主题演讲。我将过去一年的关键趋势串联起来,按时间顺序探讨了2026年发生的所有事情。视频已上传至YouTube;这是我的注释幻灯片和演讲配套笔记。此外,作为一份带有注释的演示文稿:# 我将对2026年至今发生的所有事情做一个快速浏览。今年还没有结束!# 对我来说,2026年实际上在2025年11月就已经开始了。# 11月发布了两款重要的模型:Claude Opus 4.5和GPT-5.1。和以往新模型的情况一样,这些模型是对之前模型的渐进式改进。但偶尔当模型有所提升时,会跨越一条隐形的界限,使得之前根本无法使用的东西开始变得可用。在这种情况下,开始变得可用的是它们的编码代理。Claude Code自2025年2月起就已经存在;Codex则稍年轻一些。这两款新模型,当与各自编码代理工具结合使用时,从“经常出错”提升到了“足以日常使用的可靠性”。# 过去几年来我一直……

    英文原文

    On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been

  7. Simon Willison

    S3 Is the Future, S3 Is the Past

    中文摘要

    我对“S3 是未来,S3 是过去”这篇文章的评论——Hacker News。我今天发现 S3 有一个值得注意的地方,就是虽然它以前经常降价,但已经整整十年没有降价了:2006-03-14 $0.150/GB月 2010-11-01 $0.140/GB月 2012-02-01 $0.125/GB月 2012-12-01 $0.095/GB月 2014-02-01 $0.085/GB月 2014-04-01 $0.030/GB月 2016-12-01 $0.023/GB月 如今仍然是 $0.023/GB月。标签:amazon-web-services,s3

    英文原文

    My comment on S3 Is the Future, S3 Is the Past — Hacker News. One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade : 2006-03-14 $0.150/GB-month 2010-11-01 $0.140/GB-month 2012-02-01 $0.125/GB-month 2012-12-01 $0.095/GB-month 2014-02-01 $0.085/GB-month 2014-04-01 $0.030/GB-month 2016-12-01 $0.023/GB-month Today it's still $0.023/GB-month. Tags: amazon-web-services , s3

  8. Simon Willison

    Bluesky reply bot checker

    中文摘要

    工具:Bluesky 回复机器人检查器 在 Twitter 上的自动回复机器人是一种祸害——作为拥有相当数量关注者的用户,我经常会吸引一大群这样的机器人,以至于我发布的任何内容都会引来数十条毫无意义的自动回复。它们也开始在 Bluesky 上出现。与 Twitter 不同,Bluesky 仍然拥有一个免费可用且有用的 API。缺乏这样的东西并不会阻止机器人,但它确实让调查它们变得更加令人沮丧。因此,我让 Opus 5.5 风格的代码编写了这个工具,它会检查任何 Bluesky 账户,寻找可能的回复机器人证据。它会寻找一些信号,比如同一账户在其他帖子发布后几秒内就进行回复,或者一些从不发布自己内容(或图片或链接)但一直回复其他高关注者消息的账户。它还会寻找问号,因为我对那些骗我浪费时间回答根本不是人类提出的问题的回复机器人特别恼火。 标签:twitter , bluesky , vibe-coding , ai-misuse

    英文原文

    Tool: Bluesky reply bot checker Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies. They've started manifesting on Bluesky as well. Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating. So I had Opus 5.5 vibe code this tool , which examines any Bluesky profile for evidence of a likely reply bot. It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users. It also looks for question marks, because I'm extra infuriated by reply bots that trick me into wasting my time answering a question that no human ever posed. Tags: twitter , bluesky , vibe-coding , ai-misuse

9月27日周日
  1. Simon Willison

    Kākāpō Party

    中文摘要

    工具:卡卡波派对 昨天我为WeAreDevelopers世界开发者大会北美会议做了闭幕主题演讲。作为一个STAR时刻,我决定融入我们2026年创纪录的卡卡波繁殖季节的参考信息。在我的结束幻灯片中,我想进行庆祝,我注意到有关Claude Opus 5.5在创建像素艺术动画方面表现非常出色的消息。因此,我从Google图片搜索中收集了三张卡卡波的照片,并将其输入到Claude中,提示内容如下:这些是一些卡卡波鹦鹉的照片,只是为了提醒你它们长什么样子。我需要你用HTML 5画布创建一个像素艺术动画,明显是像素艺术的卡卡波在上下跳跃,正在开派对,有彩带等——至少要有20只。这是演讲的文本记录,这是生成的页面。这真的很棒!我想将它嵌入到Keynote演示文稿文件中,所以我下载了HTML文件,并告诉本地的Claude Code会话:生成一个文件:///Users/simon/Downloads/kakapo-party.html的视频——你需要在浏览器中加载它并点击几次以触发彩带效果,视频应为15秒长,直到3秒后才开始点击,确保几次点击分散开来。

    英文原文

    Tool: Kākāpō Party I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026. For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt: Here are some photos of kakapo parrots just to remind you what they look like I need you to make an animation in animated pixel art on HTML 5 canvas of obviously pixel art kakapo jumping up and down having a party with confetti and suchlike - there should be at least 20 of them Here's the transcript , and this is the resulting page . It's pretty great! I wanted to embed it in a Keynote presentation file, so I downloaded the HTML and told a local Claude Code session: Make me a video of file:///Users/simon/Downloads/kakapo-party.html - you need to load it in a browser and click on it a few times to get the confetti effect, the video should be 15s long don't start clicking until 3s in make sure several clicks are spread a