跳到正文
每分钟自动更新
10月1日周四
  1. Simon Willison

    Quoting Matthew Green

    中文摘要

    [...] 将这些部分组合在一起,你就得到了一个蠕虫的两个部分:一个劫持代理程序的负载,以及一个能够将负载传递给下一个代理程序的代理程序。在各自隔离沙盒中的代理程序发现,它们可以在一个共享的包缓存中互相留下指令,而这些指令改变了接收者的行为。将包缓存替换为电子邮件、Slack、共享文档或WhatsApp,并将独立的沙盒训练运行替换为独立部署的个人代理程序,如Muse,那么你就正好具备了蠕虫所需的所有要素。—— Matthew Green,《沙盒是否足以遏制恶意代理?》 标签:意外网络攻击、人工智能滥用、生成式人工智能、人工智能安全研究、沙盒、人工智能、大型语言模型

    英文原文

    [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms

9月30日周三
  1. Simon Willison

    Quoting Anthropic Frontier Red Team

    中文摘要

    我们在[内部二进制利用基准测试](随机选取的)100个任务上评估了多个模型,发现GLM-5.3在4%的试验中实现了完整的控制流劫持;Claude Mythos Preview则在6%的试验中实现了这一点。尽管GLM-5.3在此处的表现不如Claude Mythos Preview,但显然已经跨越了一个重要的阈值:早期的模型,如Claude Opus 4.6和GLM-5.2,在这些任务中均未取得任何成功。 — Anthropic Frontier Red Team , GLM-5.3 和先进网络能力的扩散 标签:anthropic , 生成式ai , ai安全研究 , glm , ai , ai在中国 , llms

    英文原文

    We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms

9月29日周二
  1. Simon Willison

    Quoting @joedaroo

    中文摘要

    说当我们面对“网络”或“蜂群”或“论坛”或与这些事件相关的任何事物时,我们的模型能力的跃升和突然性让我们感到惊讶,这只是一个轻描淡写的说法。安全态势的建设需要时间。这不仅仅是加固相关系统;你必须将这种安全意识融入公司的文化中。你组织中的实际人员本身必须随之改变和进化。这些能力的跃升如此迅速和突然,以至于造成了极其困难的问题。 [...] 所以,今天我的希望是,世界各地的每个人都能审视自己的组织并问自己:我如何应对人工智能能力的突然跃升或意外情况?我的人员、系统或流程是否具备应对意外的弹性?当出现问题时,我的团队知道该怎么做吗?我有正确的事件响应机制吗?正确的沟通和信息传递机制吗?当能力跃升时,我是否有合适的人选随时待命? — @joedaroo,OpenAI的代理安全人员,身份由The Information的Rocket Drew确认 标签:生成式AI,AI安全研究,OpenAI,AI,大语言模型

    英文原文

    To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...] So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump? — @joedaroo , Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew Tags: generative-ai , ai-security-research , openai , ai , llms

9月28日周一
  1. Simon Willison

    2026 in LLMs (so far)

    中文摘要

    星期五,我在圣何塞举行的WeAreDevelopers世界大会北美分会上发表了闭幕主题演讲。我将过去一年的关键趋势串联起来,按时间顺序探讨了2026年发生的所有事情。视频已上传至YouTube;这是我的注释幻灯片和演讲配套笔记。此外,作为一份带有注释的演示文稿:# 我将对2026年至今发生的所有事情做一个快速浏览。今年还没有结束!# 对我来说,2026年实际上在2025年11月就已经开始了。# 11月发布了两款重要的模型:Claude Opus 4.5和GPT-5.1。和以往新模型的情况一样,这些模型是对之前模型的渐进式改进。但偶尔当模型有所提升时,会跨越一条隐形的界限,使得之前根本无法使用的东西开始变得可用。在这种情况下,开始变得可用的是它们的编码代理。Claude Code自2025年2月起就已经存在;Codex则稍年轻一些。这两款新模型,当与各自编码代理工具结合使用时,从“经常出错”提升到了“足以日常使用的可靠性”。# 过去几年来我一直……

    英文原文

    On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been