跳到正文
Simon Willison·· 2 天前

Qwen3.8 27B addition in words

Qwen3.8 27B addition in words

中文摘要

研究:Qwen3.8 27B 以文字形式进行加法运算的实验。Colin Frasier 在 Bluesky 上分享了他两年前使用 GPT-4o 进行的一项实验,目的是测试它在面对越来越大的数字时,能否“计算总和但以文字形式返回答案”。他分享了这些结果的图表:我确信 GPT-4o 没有作弊使用计算器,尤其是因为它在很多计算中都出错了,但我受到启发,在本地硬件(DGX Spark)上重新运行了这个实验,以在完全受控的环境中探索这一现象。我将他的图片粘贴到一个 Codex Remote 会话(GPT-6 Astra)中,并让它使用 Qwen3.8-27B-Q4_K_M.gguf 运行相同的实验。以下是每种组合进行 30 次尝试的结果,且禁用了推理功能:然后我再次运行了实验,但这次启用了推理功能。由于每对数据的处理时间更长,我并没有为每个组合运行 30 次样本,而是只运行了一次——这导致热力图的视觉效果要差很多,因为每个方块要么是 100%,要么是 0%:它在 169 次尝试中答对了 167 次,由于这些是一次性测试,我确信第二次运行会得到不同的结果。这是包含推理追踪的报告版本:

英文原文

Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%: It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here. Here's a version of the report that includes the reasoning traces from

应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。

来源:Simon Willison · simonwillison.net