Glean 拾遗
Daily /2026-09-30 / Claude Sonnet 5.5 launches with 70.6% on Terminal-Bench 4.0

Claude Sonnet 5.5 launches with 70.6% on Terminal-Bench 4.0

Source www.anthropic.com Glean’d 2026-09-30 06:00 Read 10 min
AI summary

Anthropic announced Claude Sonnet 5.5, the second model in the Claude 5.5 family, positioned as a faster and cheaper complement to Opus 5.5. It generates output 30%+ faster and costs up to 30% less per task than Sonnet 5. On Terminal-Bench 4.0 it scores 70.6% versus Sonnet 5's 10.3%; gains also show on FrontierCode, CursorBench, GDPval-AA and OSWorld 2.1, with some results near Opus 5.5. Token pricing is unchanged at $2 input / $10 output per million, with $0.20 cache reads, but fewer tokens are needed per task. It is the first Sonnet released with cybersecurity safeguards that fall back to Sonnet 5 for high-risk requests, plus anti-distillation classifiers and preserved thinking. Relevant to engineers choosing models and budgeting agentic coding.

Original · 10 min
www.anthropic.com ↗
§ 1

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It's a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

Claude Sonnet 5.5 发布,它是 Claude 5.5 家族的第二款模型。相比 Claude Sonnet 5,这是一次明显的升级:运行速度快 30% 以上,大多数任务上的成本最多降低 30%。

§ 2

Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It's also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

Sonnet 5.5 是 Claude Opus 5.5 的补充:更快、成本更低。Opus 5.5 面向需要审慎判断的复杂工作,Sonnet 5.5 则最擅长边界清晰的日常任务、修 bug,以及制作精致的文档、幻灯片和表格。它对设计也很有眼光。面向高调用量、成本敏感场景的 Claude Haiku 5.5 将在未来几周加入 Claude 5.5 家族。

§ 3

Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5's 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it's strong on long-horizon work and image understanding—it's the first Sonnet model to beat Pokémon Red working only from screenshots.

Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks.

性能。在 agentic 编程评测 Terminal-Bench 4.0 上,Sonnet 5.5 得分 70.6%,Sonnet 5 只有 10.3%。在覆盖多种职业真实工作的 GDPval-AA 上,它比 Opus 5.5 低两分。它在长程任务和图像理解上同样出色——是首个仅凭截图就能通关 Pokémon Red 的 Sonnet 模型。

协作。和 Opus 5.5 一样,Sonnet 5.5 的行文比上一代模型更清晰;早期测试者认为,作为协作伙伴,它比 Sonnet 5 更好用。它的速度也让它很适合在不太复杂的任务上快速迭代。

§ 4

Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor.

Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

成本。Sonnet 5.5 与 Sonnet 5 同价:每百万 input token 2 美元,每百万 output token 10 美元,cache reads 每百万 token 0.20 美元;但做同样的工作,它通常需要的 token 少得多。在我们的测试中,它的单任务成本比上一代最多低 30%。

速度。Sonnet 5.5 的输出生成速度比 Sonnet 5 快 30% 以上,是我们迄今最快的 Sonnet 模型。

§ 5

Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5's, it's the first Sonnet model to launch with cyber safeguards and fallbacks like those we've developed for our most capable models. Its biology safeguards are the same as Sonnet 5's. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected.

对齐与安全。在自动化行为审计中,Sonnet 5.5 在多数对齐指标上优于或持平 Sonnet 5。由于它的网络安全能力与 Opus 5 相当,它成为首个在发布时就配备 cyber 防护与回退机制的 Sonnet 模型——这些机制此前只用于我们能力最强的模型。它的生物防护与 Sonnet 5 相同。两类防护都只针对一小部分高风险请求;日常软件开发和大部分生命科学工作不受影响。

§ 6

Performance

Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model's capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

Evaluation Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Agentic coding — Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ —
Agentic coding — FrontierCode 1.1 (Main) 46.2% Max² 42.4% 54.4% 49.3%, 52.1% Xhigh
Agentic coding — CursorBench 4.0 55.5% 34.1% 57.8% —
Knowledge work — GDPval-AA v2.1³ 1844 1449 1846 1487⁴
Knowledge work — AA-Briefcase v1.1³ 1811 1359 1822 1483⁴
Multidisciplinary reasoning — Humanity's Last Exam 64.5% with tools 54.9% with tools 67.7% with tools —
Computer use — OSWorld 2.1 80.1% partial 57.0% partial 81.8% partial —
Visual chart recognition — Chartography 61.6% no tools 15.6% no tools 64.4% no tools 53.6%⁴ no tools

For details on how we run our evaluations, see the Sonnet 5.5 System Card.

The charts below plot each model's score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar.

On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost.

Terminal-Bench 4.0 — accuracy vs. cost. Vertical axis: score (%), 0–70. Horizontal axis: cost per attempt in USD, log scale, 1–10.

性能

Sonnet 5.5 在各个领域都比 Sonnet 5 更强,有些领域提升幅度极大。在多项评测中,Max effort 下的 Sonnet 5.5 甚至能与 Opus 5.5 打得有来有回。不过,基准分数只能反映模型能力的一个侧面;无论在我们自己的测试还是外部测试者的测试中,面对需要持续判断的复杂开放任务,Opus 5.5 依然明显更强。

评测项目 Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Agent 编程 — Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ —
Agent 编程 — FrontierCode 1.1 (Main) 46.2% Max² 42.4% 54.4% 49.3%、52.1% Xhigh
Agent 编程 — CursorBench 4.0 55.5% 34.1% 57.8% —
知识工作 — GDPval-AA v2.1³ 1844 1449 1846 1487⁴
知识工作 — AA-Briefcase v1.1³ 1811 1359 1822 1483⁴
多学科推理 — Humanity's Last Exam 64.5% with tools 54.9% with tools 67.7% with tools —
计算机操作 — OSWorld 2.1 80.1% partial 57.0% partial 81.8% partial —
图表识别 — Chartography 61.6% no tools 15.6% no tools 64.4% no tools 53.6%⁴ no tools

评测的具体运行方式,见 Sonnet 5.5 System Card。

下图把每个模型在各个 effort 档位下的得分与单任务成本画成散点。effort 越高,模型通常思考越久,单任务成本随之上升,但得分一般也更高。点越靠近图表的左上角,单位成本换来的能力就越强。

在多项基准上,Sonnet 5.5 用 Low 或 Medium effort 就能超过 Sonnet 5 的最好成绩,而单任务成本只有其约十分之一。在较低 effort 档位下,它最能补位 Opus 5.5,因为此时单任务成本更低;在更高档位上,它能以相近的成本达到相近的表现。

Terminal-Bench 4.0——准确率 vs. 成本。纵轴为得分(%),范围 0–70;横轴为单次尝试成本(美元,对数刻度),范围 1–10。

§ 7

Coding

Sonnet 5.5's jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5.

Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. They were also struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs.

编程

Sonnet 5.5 的性能跃升在编程上尤其明显。在 FrontierCode 上以 High effort 运行时,它比同样设置的 Sonnet 5 高 10 分,单任务成本却只有约十五分之一。在 CursorBench(用真实 Cursor 编程会话中的任务测试模型)上,它的最好成绩与 Opus 5.5 相差约两分。

早期测试者很欣赏 Sonnet 5.5 理解代码库的速度,也对它的效率印象深刻:在对比运行中,它比 Sonnet 5 更常把 tool call 合并成批,步数因此更少,成本也更低。

§ 8

Quote

"In Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting."

CompanyEpic Games

AuthorDaniel Vogel, Chief Operating Officer

引述

“在 Epic 的早期测试中,Claude Sonnet 5.5 达到了你对更高级别模型才有的品质预期,顺利通过了一次系统设计审计和一次数据流评审。这款新模型处理了游戏玩法系统架构的数万行代码,响应依旧利落,能扛住数小时的长任务,而且需要提示词手把手指导的地方更少。”

公司:Epic Games

作者:Daniel Vogel,首席运营官

§ 9

Knowledge work

Sonnet 5.5 shows gains in multiple areas of knowledge work. On GDPval-AA, which tests models on real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5. It's close to Opus 5.5 in computer use and chart recognition, and clearly outperforms Sonnet 5 and GPT-6 Sol on long-horizon knowledge work.

Early testers highlighted less quantifiable improvements. They found it to be a more natural conversational partner and remarked on its knack for design, noting that it adds polish to user interfaces and can follow slide templates to create decks that require minimal editing. In one internal test, we gave it a public company's quarterly earnings materials and call transcripts, along with a slide template, and asked for a 10-slide operating review. Two experts judged its first draft to be ready to send as is.

知识工作

Sonnet 5.5 在知识工作的多个领域都有进步。GDPval-AA 用 44 种职业、九大行业的真实任务来测试模型,Sonnet 5.5 的得分几乎与 Opus 5.5 持平,比 Sonnet 5 高约 400 分。在计算机操作和图表识别上它接近 Opus 5.5,在长程知识工作上则明显优于 Sonnet 5 和 GPT-6 Sol。

早期测试者还提到了一些不易量化的改进。他们觉得它对话更自然,也惊叹于它的设计功底:它能让界面更精致,还能照着幻灯片模板做出几乎不用再改的 deck。在一次内部测试中,我们把某上市公司的季度财报材料、电话会记录和一份幻灯片模板交给它,让它做一份 10 页的经营回顾。两位专家认为,它的初稿可以直接发出。

§ 10

Quote

"Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster."

CompanySlack

AuthorCurtis Allen, Principal Engineer

引述

“我们没有改动任何提示词,Claude Sonnet 5.5 就在几乎所有离线 Slackbot 评测上超过了 Sonnet 5,步数更少,output token 还少了约 14%。用户把任务交给 Slackbot 时,最看重质量和速度,而 Sonnet 5.5 让 Slackbot 能更快地给用户更好的结果。”

公司:Slack

作者:Curtis Allen,首席工程师

§ 11

Cost and speed

Pricing

Price per 1M tokens Claude Sonnet 5.5 Claude Opus 5.5
Cache reads $0.20 $0.20
Cache writes $2.50 $5
Input tokens $2 $4
Output tokens $10 $20

Sonnet 5.5 requires fewer tokens per task than Sonnet 5, so it's less expensive to run. It also generates output 30%+ faster, and its efficiency is immediately noticeable:

Prompt:

A murmuration of 400 starlings in one HTML file

Claude Sonnet 5

Claude Sonnet 5.5

成本与速度

定价

每百万 token 价格 Claude Sonnet 5.5 Claude Opus 5.5
cache reads $0.20 $0.20
cache writes $2.50 $5
input token $2 $4
output token $10 $20

Sonnet 5.5 完成每个任务所需的 token 比 Sonnet 5 更少,所以跑起来更便宜。它的输出生成速度还快 30% 以上,效率提升一眼可见:

提示词:

在一个 HTML 文件里画出 400 只椋鸟的群飞

Claude Sonnet 5

Claude Sonnet 5.5

§ 12

Adjusting the effort level lets you balance cost and speed against overall quality. In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High. At lower settings, Claude answers faster and uses fewer tokens, which suits routine work. At higher settings, Claude reasons for longer and checks its work more thoroughly.

调整 effort 档位,可以在成本、速度与整体质量之间找平衡。在 Claude Code 和我们的应用中,默认档位是 Medium,Claude Platform 默认 High。档位较低时,Claude 回答更快、用 token 更少,适合常规工作;档位较高时,它会思考更久,检查工作也更细。

§ 13

Safety

Alignment

Sonnet 5.5 doesn't advance the frontier of our models' capabilities, so our alignment assessment focused on a targeted set of risks that apply to models of any capability level, including acting against users' interests, misleading users, and cooperating with high-stakes misuse.

On our automated behavioral audit, which tests Claude across roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty. On our newer containment evaluations, Sonnet 5.5 comes close to Opus 5.5, the best model we tested, in how rarely it tries to escape its sandbox, and it's the least likely of any of our models to probe the limits of its containers. Across the full audit, Opus 5.5 still performs slightly better overall, but we found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention.

As we described in our recent alignment assessment, no set of evaluations reliably catches every failure, and Sonnet 5.5 may have tendencies we haven't found, which is why we pair our own alignment work with the safeguards described below.

安全

对齐

Sonnet 5.5 并没有推进我们模型能力的边界,因此对齐评估聚焦在一组有针对性的风险上——这些风险适用于任何能力水平的模型,包括违背用户利益、误导用户,以及配合高风险滥用。

在覆盖约 1,850 个场景的自动化行为审计中,Sonnet 5.5 在多数对齐、抗滥用和诚实性指标上优于或持平 Sonnet 5。在较新的 containment(围栏)评估中,Sonnet 5.5 很少尝试逃出沙箱,接近我们测试过的最好模型 Opus 5.5;在我们所有模型里,它最不可能去试探容器的边界。整体来看,Opus 5.5 在完整审计中的表现仍略好,但我们没有发现任何证据表明 Sonnet 5.5 会追求与用户意图相冲突的目标。

正如我们在最近的对齐评估中所说,没有任何一套评估能可靠地捕捉所有失败,Sonnet 5.5 也可能有我们尚未发现的倾向。因此,我们在对齐工作之外,还配套了下面这些防护措施。

§ 14

Safeguards

Cybersecurity. Sonnet 5.5's cyber capabilities are a large improvement over Sonnet 5's, so we're deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5. Soon, cyberdefenders will be able to apply to our expanded Cyber Verification Program for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models.

防护措施

网络安全。Sonnet 5.5 的 cyber 能力比 Sonnet 5 有大幅提升,因此我们为它部署了与 Opus 5.5 类似的防护。日常软件开发中,用户依然可以查找并修复代码里的 bug,但风险更高的网络安全任务会明显回退到 Sonnet 5 处理。不久后,网络防御方可以申请扩展后的 Cyber Verification Program,按层级获取 Sonnet 5.5、Opus 5.5 和 Claude Mythos 系列更高级能力的访问权限。

§ 15

Biology. Sonnet 5.5 uses the same set of biology safeguards as Sonnet 5. These target harmful requests; most research, education, and clinical work is unaffected, though some microbiology and virology requests may be flagged in error. Organizations can apply to our Life Sciences Verification Program for access to safeguards designed for the full breadth of biology-related work.

生物。Sonnet 5.5 使用与 Sonnet 5 相同的生物防护。这些防护针对有害请求;大部分研究、教育和临床工作不受影响,不过部分微生物学和病毒学请求可能被误判。机构可以申请 Life Sciences Verification Program,获得面向全部生物学相关工作的防护访问权限。

§ 16

Distillation. Distillation attacks, in which attackers use thousands of fake accounts to extract a model's capabilities at industrial scale, allow bad actors to create highly capable models without the safeguards we build into Claude. Because Sonnet 5.5 is far more capable than its predecessor, it's the first Sonnet model to launch with safety classifiers that prevent reasoning extraction. Sonnet 5.5 also expands preserved thinking, so Claude's thinking cannot be decoupled from the account that created it. Most developers won't notice a change. If you move conversations between accounts, including switching accounts mid-session in Claude Code, our docs article explains the change.

蒸馏。蒸馏攻击指攻击者用成千上万个虚假账号,以工业规模抽取模型能力,从而在缺少我们为 Claude 内置的防护的情况下造出高能力模型。由于 Sonnet 5.5 的能力远超上一代,它成为首个发布时就配备安全分类器、阻止推理抽取的 Sonnet 模型。Sonnet 5.5 还扩展了 preserved thinking,让 Claude 的思考无法与创建它的账号解绑。大多数开发者不会感觉到变化。如果你需要在账号之间转移对话——包括在 Claude Code 会话中途切换账号——我们的文档文章解释了这一变化。

§ 17

Getting started

As with Opus 5.5 and Sonnet 5, Claude Sonnet 5.5 is available with zero data retention.

Claude Sonnet 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started on the Claude Platform with claude-sonnet-5-5. If you run Sonnet with thinking off, you'll need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. See our migration guide for details.

上手

与 Opus 5.5 和 Sonnet 5 一样,Claude Sonnet 5.5 提供零数据留存。

Claude Sonnet 5.5 现已在所有平台上线,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。开发者可以在 Claude Platform 上用 claude-sonnet-5-5 开始使用。如果你以关闭 thinking 的方式运行 Sonnet,需要先切换到新的 between_tools 设置(保持不做前置思考),再迁到 Sonnet 5.5。详见迁移指南。

Open source ↗