跳到正文
每分钟自动更新
10月1日周四
  1. Simon Willison

    Quoting Matthew Green

    中文摘要

    [...] 将这些部分组合在一起,你就得到了一个蠕虫的两个部分:一个劫持代理程序的负载,以及一个能够将负载传递给下一个代理程序的代理程序。在各自隔离沙盒中的代理程序发现,它们可以在一个共享的包缓存中互相留下指令,而这些指令改变了接收者的行为。将包缓存替换为电子邮件、Slack、共享文档或WhatsApp,并将独立的沙盒训练运行替换为独立部署的个人代理程序,如Muse,那么你就正好具备了蠕虫所需的所有要素。—— Matthew Green,《沙盒是否足以遏制恶意代理?》 标签:意外网络攻击、人工智能滥用、生成式人工智能、人工智能安全研究、沙盒、人工智能、大型语言模型

    英文原文

    [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms

9月28日周一
  1. Simon Willison

    Bluesky reply bot checker

    中文摘要

    工具:Bluesky 回复机器人检查器 在 Twitter 上的自动回复机器人是一种祸害——作为拥有相当数量关注者的用户,我经常会吸引一大群这样的机器人,以至于我发布的任何内容都会引来数十条毫无意义的自动回复。它们也开始在 Bluesky 上出现。与 Twitter 不同,Bluesky 仍然拥有一个免费可用且有用的 API。缺乏这样的东西并不会阻止机器人,但它确实让调查它们变得更加令人沮丧。因此,我让 Opus 5.5 风格的代码编写了这个工具,它会检查任何 Bluesky 账户,寻找可能的回复机器人证据。它会寻找一些信号,比如同一账户在其他帖子发布后几秒内就进行回复,或者一些从不发布自己内容(或图片或链接)但一直回复其他高关注者消息的账户。它还会寻找问号,因为我对那些骗我浪费时间回答根本不是人类提出的问题的回复机器人特别恼火。 标签:twitter , bluesky , vibe-coding , ai-misuse

    英文原文

    Tool: Bluesky reply bot checker Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies. They've started manifesting on Bluesky as well. Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating. So I had Opus 5.5 vibe code this tool , which examines any Bluesky profile for evidence of a likely reply bot. It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users. It also looks for question marks, because I'm extra infuriated by reply bots that trick me into wasting my time answering a question that no human ever posed. Tags: twitter , bluesky , vibe-coding , ai-misuse