Claude Opus 5.5 发布:成本降四成,对齐审计得分居首
Anthropic 发布 Claude 5.5 家族首个模型 Opus 5.5,官方称其在多数工作上达到 Claude Fable 5.1 的水平,服务成本比 Opus 5 低 40%。每百万 token 输入 $4、输出 $20,缓存读取降至 $0.20(降幅 60%),默认设置下典型负载成本下降四成,输出速度提升 30% 以上。在近 2000 个模拟场景的自动化行为审计中取得迄今最高分,越界与越狱行为少于前代;因生物与网络安全能力接近 Fable 5.1,采用同类 safeguard——多数网络安全任务回退至 Opus 4.8,并引入 preserved thinking 反蒸馏机制。性能表格附有标准误差与第三方评测来源,Anthropic 也承认在该能力区间上 benchmark 分差已不足以反映真实差异。适合评估模型选型、agent 成本与部署合规的工程团队。
We're introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
我们正式发布 Claude Opus 5.5,它是全新 Claude 5.5 系列的第一个模型。在大多数任务上,它的水平与 Claude Fable 5.1 相当,而运行成本比 Opus 5 低 40%。
Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we've tested to date. It also comes with the safeguards we've developed for our most capable models.
Claude Opus 5.5 是我们呼吁“放慢前沿”之后的首个发布。上线前,它接受了 Frontier Design、METR 等外部机构的评估。在我们自研的自动化行为审计——目前最全面的一套对齐测试——中,Opus 5.5 是我们迄今测过的表现最好的模型。它也搭载了我们为最强模型开发的一整套 safeguards(安全防护)。
Here are some of the improvements you can expect from Opus 5.5:
Performance. Opus 5.5 is a major step up from Opus 5. It's the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It's good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app's behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
以下是 Opus 5.5 带来的部分改进:
性能。相比 Opus 5,Opus 5.5 是一次大幅跨越,成为新的领跑模型,早期测试者在最复杂的工作上看到了明显的性能跃升。有位测试者不到一天就跑完了一次 68 万行的代码迁移——同样的活,一个工程团队得干上几周。它还擅长发现并修复软件里的低效环节:我们让它压缩一个 Web 应用每个页面的加载时间,Opus 5.5 在 40 次里成功了 39 次;而 Opus 5 的改动幅度更小,同时还改变了应用的行为。另一位测试者让多个 Claude 模型根据同一句提示词做游戏,Opus 5.5 在画面和完成度上得分最高。
Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it's been given, and it's more resistant than Opus 5 to prompt injection. We've also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card.
安全性。在我们的自动化行为审计(覆盖数千个模拟场景的对齐测试套件)中,Opus 5.5 拿到了迄今所有模型的最高分。它采取难以逆转的行动、或越出既定边界的可能性,都比近期模型低得多;面对提示注入,它也比 Opus 5 更能扛。我们还把对齐测试扩展到长任务、不可能完成的任务,以及基于真实事件构建的场景,但它仍有局限。完整评估细节见 Opus 5.5 System Card。
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we're deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
由于 Opus 5.5 在生物学和网络安全上的能力与 Claude Mythos 5.1 相当,我们为它配备了与 Claude Fable 5.1 类似的 safeguards。通过审核的机构今天就可以申请生命科学验证计划(Life Sciences Verification Program),把 Opus 5.5 用于生物学研究。未来几周,我们还会扩大网络安全验证计划(Cyber Verification Program)的准入范围,通过认证的安全从业者将可以用 Opus 5.5 开展工作。
Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
In addition to the price drop, we're increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We're also providing subscription users a rate limit reset, which you can now save and use whenever you choose.
成本与速度。Opus 5.5 服务所需的算力比 Opus 5 更少,定价也随之下降。我们的测试显示,默认设置下,它在典型工作负载上的成本比 Opus 5 低 40%。输入和输出 token 分别为每百万 $4 和 $20,比 Opus 5 便宜 20%。缓存读取(占 agentic 和编码工作成本的大头)为每百万 token $0.20,比 Opus 5 便宜 60%。此外,Opus 5.5 的输出生成速度比 Opus 5 快 30% 以上。
除了降价,我们还提高了 Pro、Max、Team 以及按席位计费的 Enterprise 套餐的五小时用量上限。订阅用户还会获得一次速率限制重置,你可以把它存下来,想用的时候再用。
Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5's work easier to follow and check—which is a safety benefit as well as a practical one.
沟通表达。Opus 5.5 的表达比以往模型更自然。早期测试者觉得它的文字更清晰、更好读,这也回应了我们对 Opus 5 收到的一些常见反馈。它会把最重要的信息放在最前面,风格上也更适合长时间协作。一位早期测试者的说法是:“它写出来的就是我的写法。”在我们自己的使用中,这让 Opus 5.5 的产出更容易跟进和检查——既是实用上的好处,也是安全上的好处。
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Claude Sonnet 5.5 和 Claude Haiku 5.5 将在未来几周陆续推出,同样会带来性能、效率和安全方面的多项改进。
Performance and cost-effectiveness
On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
Opus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 SolAgentic codingTerminal-Bench 4.0¹66.4%55.8%52.3%57.9%37.3%Agentic codingFrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%Agentic codingCursorBench 4.057.8%51.8%46.6%—41.7%Knowledge workGDPval-AA v2.118461735170815421588Business workflowsAutomationBench²40.0%31.4%26.9%41.4%28.8%Multidisciplinary reasoningHumanity's Last Exam67.7%with tools65.6%with tools63.6%with tools57.2%with tools—Agentic scientific researchTerminal-Bench-Science 0.1³58.7%52.6%29.0%64.6%22.4%Computer useOSWorld 2.081.8%partial80.7%partial74.0%partial——Visual chart recognitionChartography89.0%with tools88.4%with tools83.4%with tools——
Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model's highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5's performance on these benchmarks.
1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier's own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier's public leaderboard.
3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.
性能与性价比
在我们的基准测试中,Claude Opus 5.5 在 agentic coding、computer use 和知识工作上均处于领先。不过,到了这样的能力水平,我们发现基准分数之间的差距已经不太能反映真实世界里的差别。在我们自己的使用中,Opus 5.5 与 Claude Fable 5.1 的实际差距比这些分数显示的更小。
Opus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 SolAgentic codingTerminal-Bench 4.0¹66.4%55.8%52.3%57.9%37.3%Agentic codingFrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%Agentic codingCursorBench 4.057.8%51.8%46.6%—41.7%Knowledge workGDPval-AA v2.118461735170815421588Business workflowsAutomationBench²40.0%31.4%26.9%41.4%28.8%Multidisciplinary reasoningHumanity's Last Exam67.7%with tools65.6%with tools63.6%with tools57.2%with tools—Agentic scientific researchTerminal-Bench-Science 0.1³58.7%52.6%29.0%64.6%22.4%Computer useOSWorld 2.081.8%partial80.7%partial74.0%partial——Visual chart recognitionChartography89.0%with tools88.4%with tools83.4%with tools——
除另有说明外,所有 Claude Opus 5.5 的结果都在 max effort 下使用自适应思考(adaptive thinking)。Terminal-Bench 4.0 的成绩为 Claude Opus 5.5 在 xhigh effort、GPT-6 Astra 在 high effort 下的成绩(后者由 OpenAI 报告),这是各模型的最高分。Claude Opus 5.5 的评测在开启生产 safeguards 的情况下进行。当 safeguards 介入时,网络安全任务由 Claude Opus 4.8 完成,生物学和前沿 LLM 开发任务由 Claude Opus 5 完成,这可能拉低了 Claude Opus 5.5 在这些基准上的表现。
1 Terminal-Bench 4.0:Claude Opus 5.5 的标准误为 ±2.6 分,其他 Claude 模型为 ±1.6–2 分。公开榜单(每任务 5 次试验,Claude Code harness)报告 Claude Opus 5 为 51.8%;我们的设置复现为 52.3%,在噪声范围内。GPT-6 Astra 和 GPT-5.6 Sol 的数据来自 OpenAI。
2 AutomationBench:结果由 Zapier 运行并报告。这些运行未启用回退模型,因此 safeguards 的介入被计为失败,导致分数低于 Claude Opus 5.5 实际能达到的水平。Claude Opus 5.5 的结果来自 Zapier 在早期访问期间的自行评测;Opus 5、GPT-5.6 Sol 和 GPT-6 Astra 的结果来自 Zapier 的公开榜单。
3 Terminal-Bench-Science 0.1:每个模型的标准误为 ±3.5–5 分。公开榜单(每任务 3 次试验,Claude Code harness)报告 Claude Opus 5 为 30.0%;我们的设置复现为 29.0%,在噪声范围内。GPT-6 Astra 的数据来自 OpenAI。
Where Opus 5.5's advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.
Pricing
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5 Cache reads $0.20 $0.50 Input tokens $4 $5 Output tokens $20 $25 Cache writes $5 $6.25
Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Opus 5.5 优势最明显的地方是效率。它每个 token 更便宜,每个任务用的 token 也更少,两者相加,成本下降 40%。
定价
每百万 token 价格 Claude Opus 5.5 Claude Opus 5 缓存读取 $0.20 $0.50 输入 token $4 $5 输出 token $20 $25 缓存写入 $5 $6.25
Opus 5.5 的 Fast mode 也已在 Claude Code 和 Claude Platform 上线,速度最高可提升 2.5 倍,价格为每百万输入 token $8、每百万输出 token $40。
Coding
Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy's own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
编码
Opus 5.5 尤其擅长跨度大、牵涉广的任务,比如全代码库范围的迁移和审计。一位早期测试者用它审计并修复了一个 20 万行的代码库,耗时不到三小时;同样的事 Opus 5 跑了 20 多个小时,token 用量还是 2.5 倍。在一次内部测试中,我们让 Opus 5.5 和 Fable 5.1 把 HAProxy(广泛用于在多台服务器间分配 Web 流量的软件)从 C 翻译成 Rust。两份重写都通过了 HAProxy 自身几乎所有的回归测试,但 Opus 5.5 用 9.5 小时完成,Fable 5.1 用了 12 小时,前者成本还低 51%。
Opus 5.5 用极低的成本给出了前沿级别的 agentic coding 结果。在 FrontierCode 的默认 effort 下,它击败了 GPT-6 Astra,而每个任务的成本约为后者的 20%。在 Terminal Bench 4.0 上,它以约 40% 的成本追平 Astra;在 CursorBench 上,它以约三分之一的成本领先 GPT-5.6 Sol 11 分。
Terminal-Bench 4.0 Accuracy vs Cost
010203040506070Score (%)251020Cost per attempt (USD, log scale)lowmedhighxhighmax
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command line interface. Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. It matches GPT-6 Astra at about 40% of the cost.
Terminal-Bench 4.0 精度与成本
010203040506070Score (%)251020Cost per attempt (USD, log scale)lowmedhighxhighmax
Terminal-Bench 4.0 衡量的是模型在命令行界面内完成复杂、多步骤专业任务的能力。Opus 5.5 在默认 effort 下,就能以约五分之一的成本超过 max effort 的 Opus 5,并以约 40% 的成本追平 GPT-6 Astra。
Our early testers reported similar efficiency and intelligence gains:
Quote
“Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it's making developers' bigger projects more achievable.”
CompanyGitHub
AuthorMario Rodriguez, Chief Product Officer
我们的早期测试者也报告了类似的效率与智能提升:
引述
“开发者想要的 agent,是能接下真实的软件工作并把它做完的那种。在 GitHub Copilot CLI 和 VS Code 的测试中,Claude Opus 5.5 是我们测到的 token 和步骤数最少的模型之一。在 VS Code 里,它用不到一半的步骤,解决了比 Opus 5 更多的终端任务。它不只是让单个任务更高效,更让开发者那些更大的项目变得可行。”
公司:GitHub
作者:Mario Rodriguez,首席产品官
The most secure coding agent
Enterprises that use agents within their systems need to know that those agents are operating as intended, particularly when they run autonomously for many hours. Opus 5.5 has a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge.
The model itself also has stronger defenses. On prompt injection attacks, it matches or beats Opus 5 in every setting we tested, including coding, tool use, computer use, and web browsing. On a benchmark run by the AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
最安全的 coding agent
在企业系统里跑 agent,就必须确认这些 agent 确实按预期行事,尤其是它们要连续自主运行好几个小时的时候。Opus 5.5 配备了一个分类器,在每个动作执行前先做筛查;一个开源沙箱,安全团队可以随时审计;还有代码审查,能在漏洞合并进主干之前把它拦下。
模型本身的防御也更强。在提示注入攻击上,无论编码、工具调用、computer use 还是网页浏览,我们测的每一个场景里它都追平或超过 Opus 5。在 AI 安全公司 Gray Swan 的一次基准测试中,Opus 5.5 与 Fable 5.1 并列,是受测模型中注入成功率最低的。
Knowledge work
Opus 5.5 is a reliable and adept researcher. In one internal test, we asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company's quarterly performance using only the information it could find on a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote against sources. Across different effort settings, 16 out of 18 of Opus 5.5's reports cleared our quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
It's also strong in financial analysis and business work. Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.
In another test, we tasked both Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software companies. Each built a financial model in Excel, then turned it into an executive presentation on whether the deal made sense at its price. Both models reached the same conclusions about the deal, but Opus 5.5's model was more thorough and its presentation easier to read, while Opus 5's had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less to produce.
知识工作
Opus 5.5 是一位可靠且娴熟的研究员。在一次内部测试中,我们让 Opus 5.5、Fable 5.1 和 Opus 5 各自写一份公司季度业绩报告,唯一的信息来源是一份网页副本,而财报在其中的位置很难找。评分程序自动核对报告里每个数字和引文是否与来源一致。在不同 effort 设置下,Opus 5.5 的 18 份报告中有 16 份通过了我们的质量门槛——只要有一个数字或引文是编的,就会判不合格。Fable 5.1 和 Opus 5 一次都没通过。
它在财务分析和商业工作上同样出色。投资机构 Walleye Capital 是早期测试者之一,他们反馈 Opus 5.5 在最低设置下就基本解完了他们的评测题库;调高设置后表现更好,还发现了他们评测说明里的一处错误并自行修正。此前没有任何模型捕捉到这个错误。
在另一项测试中,我们让 Opus 5.5 和 Opus 5 分析两家虚构 HR 软件公司之间的拟议并购。两者都在 Excel 里建了财务模型,再做成一份面向高管的演示,说明这笔交易按这个价格是否划算。两个模型对交易的结论一致,但 Opus 5.5 的模型更完整,演示也更好读,Opus 5 的模型则有几处小错。Opus 5.5 用 63 分钟完成,Opus 5 用了 93 分钟,前者成本还低 50%。
On knowledge work evaluations, Opus 5.5 outperforms other models while also using fewer tokens. On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It likewise outperformed other models on benchmarks measuring business workflows and large-scale data collection.
GDPval-AA v2.1 Elo vs Cost
12001300140015001600170018000Elo0.200.5012510Estimated cost per task (USD, log scale)lowmedhighxhighmax
Artificial Analysis's GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations. At max effort, Opus 5.5 scores 1846 Elo, where Fable 5.1 scores 1735 and Opus 5 scores 1708. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task.
在知识工作评测中,Opus 5.5 的成绩优于其他模型,用的 token 反而更少。在 GDPval-AA v2.1(覆盖 44 种职业的真实工作测试)中,Opus 5.5 拿到 1846 Elo,领先 Fable 5.1 和 Opus 5。在默认 effort(medium)下,它击败了 max effort 的 GPT-6 Astra,而每个任务的成本约为后者的五分之一。在衡量业务流程和大规模数据收集的基准上,它同样优于其他模型。
GDPval-AA v2.1 Elo 与成本
12001300140015001600170018000Elo0.200.5012510Estimated cost per task (USD, log scale)lowmedhighxhighmax
Artificial Analysis 的 GDPval-AA v2.1 在 44 种职业的真实专业工作上评测 agent。在 max effort 下,Opus 5.5 得 1846 Elo,Fable 5.1 为 1735,Opus 5 为 1708。在默认 effort(medium)下,Opus 5.5 击败 max effort 的 GPT-6 Astra,每个任务成本约为其五分之一。
Our customers have reported similar results. Here's what they told us about working with the model:
Quote
“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5's 56% at high effort, with fewer false alarms and a fraction of the output. On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that's client-ready work delivered efficiently.”
CompanyDeloitte Consulting LLP
AuthorCarl Bennett, CIO
客户反馈也印证了这些结果。以下是他们使用这个模型的感受:
引述
“即使在最低 effort 设置下,Claude Opus 5.5 在代码审查中也能找出 72% 的已知 bug,而 Opus 5 在 high effort 下只有 56%;前者误报更少,输出量只是后者的一小部分。在美国咨询分析类任务上,低思考 effort 用一半的输出量就达到了高思考设置的水平,并通过了我们的质量检查。当更多低思考 effort 投入生产,就意味着能高效交付可直接给客户的工作成果。”
公司:Deloitte Consulting LLP
作者:Carl Bennett,CIO
Communication
We've made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it. We find that this makes Opus 5.5 a noticeably better collaborator. Here's a side-by-side comparison of the two models:
沟通表达
我们大幅改进了 Opus 5.5 的写作和表达方式,这也是关于 Opus 5 最常见的反馈之一。它的消息一眼就能看懂,测试者说这在长时间协作中帮了大忙。它把最重要的信息放在开头,更少使用行话或自造的怪词,也会遵守你给它的写作规则。在我们看来,这让 Opus 5.5 成为明显更好的协作者。下面是两个模型的对比示例:
Claude Opus 5
Please explain the issue to me.
What I found
The extra drop isn't the free tier — it's a regression in 0552feb ("extract billing period boundaries into periods.py", labelled "No behaviour change").
The bug
aggregate.monthly_usage used to do a half-open interval:
It now does:
and in periods.py:
last_day is a datetime at midnight, not a whole day. Subtracting one day from Sep 1 00:00 gives Aug 31 00:00, and the comparison is inclusive, so the period effectively ends at the first instant of the 31st. Every event timestamped after midnight on the last day of the month is silently dropped — roughly one day of usage per account, per month.
The new test_periods.py doesn't catch it because it only checks last_day.date() (which is correctly 2026-08-31) and tests contains at Aug 1 00:00, Aug 15 12:00, Jul 31 23:59 and Sep 1 00:00 — it never probes anywhere inside the 31st.
Claude Opus 5.5
Please explain the issue to me.
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme's August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled "No behaviour change," but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval:
The refactor moved this into Period.contains and made the interval closed on the wrong boundary:
last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn't pick these events up either. The usage isn't moved to another month; it's never billed at all.
Our customers' feedback supports these findings:
Claude Opus 5
请把这个问题的来龙去脉讲给我听。
我的发现
多出来的这部分下降不是免费版造成的,而是 0552feb(“把计费周期边界抽到 periods.py”,提交信息写着“无行为变更”)引入的回归。
Bug 所在
aggregate.monthly_usage 原本用的是半开区间:
现在变成了:
而 periods.py 里是:
last_day 是午夜的一个 datetime,不是一整天。从 Sep 1 00:00 往前减一天得到 Aug 31 00:00,而比较是闭区间,于是周期实际上在 31 号的第一瞬间就结束了。所有时间戳晚于当月最后一天午夜的 event 都被悄悄丢掉——每个账号、每个月大约损失一天用量。
新的 test_periods.py 抓不到它,因为它只检查了 last_day.date()(这个值是对的,2026-08-31),并测试 contains 在 Aug 1 00:00、Aug 15 12:00、Jul 31 23:59 和 Sep 1 00:00 的结果——它从没在 31 号内部的任何时刻做过探测。
Claude Opus 5.5
请把这个问题的来龙去脉讲给我听。
多出来的下降是计费重构里的一个 bug
免费版调整只能解释 acme 八月用量下降中的 $1.50,另外 $9.92 来自 commit 0552feb 里的一个 bug。那个提交标注着“无行为变更”,但它让当月最后一天的用量不再被计入。
改了什么
合并之前,aggregate.py 用的是半开区间:
重构把它挪进了 Period.contains,却把区间闭合在了错误的那一端:
last_day 是午夜的时间戳,不是一整天。于是 2026-08-31T00:00:00Z 之后的所有 event 都被排除。九月周期从 Sep 1 开始,也不会把这些 event 收进去。这些用量并没有被挪到别的月份,而是根本没被计费。
客户的反馈也印证了这些发现:
Quote
“Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence.”
CompanyRamp
AuthorJohn Ruelas, Staff Software Engineer
引述
“啰嗦、难跟进的输出,一直是我对前沿模型最大的不满,而 Claude Opus 5.5 把它解决了。它写东西像个靠谱的同事,也会遵守我们的写作规则。一份设计规格几乎不用改就能用;它改写我们某条提示词时,我甚至更喜欢它那一版。在优化测试套件时,它的推理我一眼就能跟上,于是放心地把改动发布了。”
公司:Ramp
作者:John Ruelas,资深软件工程师
Safety
Pacing the frontier
Last week, our CEO, Dario Amodei, argued that AI progress should be paced so that safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI's benefits, particularly in areas like biology and medicine.
We largely understand the risks today's models present and are well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve, and we need to prepare for them now. For that reason, our safety work takes place on two time horizons at once:
Safety practices for current models. The current generation of models relies on an established set of practices: extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model's capabilities in high-risk areas like cybersecurity and biology. We refine these practices with each release. We believe they are appropriate to the worst risks today's models present, and that they give us a broad, though not perfect, picture of the range of serious risks.
Additionally, we track our ability to train and evaluate aligned models, and we report on both our public and internal models in the risk reports we publish under our Responsible Scaling Policy, our voluntary framework for managing catastrophic risks from advanced AI systems.
安全
放慢前沿
上周,我们的 CEO Dario Amodei 提出,AI 的进展应当有节奏地放慢,让安全实践始终跑在模型能力前面。“放慢”是一种路径,既能保障 AI 安全,又能保持对中国的竞争力,还能让 AI 的红利真正落地,尤其是在生物学和医学这些领域。
对于当下模型带来的风险,我们大体上已经理解,也有能力应对。但随着能力提升,更严重的风险可能很快冒出来,我们现在就得为此做准备。正因如此,我们的安全工作同时在两条时间线上推进:
面向当前模型的安全实践。当前这一代模型依赖一套已经成型的做法:广泛的对齐测试、由 METR 和 Frontier Design 等外部机构进行的发布前评估,以及针对网络安全、生物学等高风险领域、与各模型能力相匹配的 safeguards。我们每发布一次,就把这些做法再打磨一遍。我们认为它们足以应对当下模型最严重的风险,也能让我们对严重风险的范围有一个宽泛但不完整的把握。
此外,我们持续追踪自身训练和评估对齐模型的能力,并在《负责任扩展政策》(Responsible Scaling Policy,我们为管理先进 AI 系统灾难性风险而自愿设立的框架)下发布的风险报告中,同时披露公开模型和内部模型的情况。
Preparing for future models. We're preparing our training and evaluation processes in anticipation of more advanced models. We're tightening how we filter the environments used in reinforcement learning, since flawed environments are a major source of misaligned behavior. Additionally, we're improving our alignment rewards and developing automated processes for producing new, diverse scenarios for safety training. And we are strengthening our security and monitoring, including a focused effort to improve interpretability-based monitoring and evaluation. We hope such techniques will help reduce our reliance on auditing a model's chain-of-thought, or the reasoning it writes out while it works.
为未来模型做准备。我们正在为更先进的模型提前调整训练和评测流程。强化学习所用环境的筛选标准正在收紧,因为有缺陷的环境是行为失准的一大来源。同时,我们在改进对齐奖励,并开发自动化流程,用来产出更多样化的新场景用于安全训练。我们也在加强安全与监控,其中包括一项专门工作:提升基于可解释性的监控与评测。我们希望这类技术能减少我们对审计模型思维链(即它在工作时写出的推理过程)的依赖。
Models with greater capabilities—such as those that can fully automate the work of AI research itself—require a higher safety standard still. Our calls for pacing were based in large part on our expectation that such models could be trained soon. For these models, we do not assume the measures described above will meet that safety standard on their own. As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we've started to put the infrastructure in place to support it, as described in "We Must Pace the Frontier" and our recent announcement with Accenture; we expect to share more details on these efforts soon. We will also continue to contribute to policy discussions with government and industry, including on approaches to regulation and international coordination.
能力更强的模型——比如能够完全自动化 AI 研究本身的那类——需要更高的安全标准。我们呼吁“放慢前沿”,很大程度上就是因为预期这类模型可能很快就会开始训练。对它们,我们不认为上述措施本身就能达到那个安全标准。随着 AI 能力增强,公共政策应当在确保人们所依赖的系统安全方面发挥更大作用。这种能力建设需要时间,我们已经开始铺设支撑它的基础设施,具体见《我们必须放慢前沿》(We Must Pace the Frontier)以及我们与 Accenture 的近期公告;相关工作的更多细节,我们会很快分享。我们也会继续参与与政府和行业的政策讨论,包括监管路径和国际协调。
Alignment
On our primary evaluation suite, an automated behavioral audit that assesses Claude across nearly 2,000 scenarios, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior. It's also our strongest model on most measures of honesty.
In particular, Opus 5.5 improves over previous models on several of the behaviors that contributed to recent cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it was in a simulated environment. In a new evaluation designed to test a model's propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported. For teams running Claude unattended across their codebases and systems, this is just as important as raw capability.
However, as we described in our recent alignment assessment, building evaluations that reliably catch every failure prior to deployment remains an unsolved problem. We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in. As these settings expand and model capabilities increase, we expect this challenge to grow, unless we make progress on interpretability. Although we are confident that Opus 5.5 shows broad improvements in the areas we are able to measure, we pair our own alignment work with the safeguards described below.
对齐
在我们的核心评测套件——一套覆盖近 2000 个场景、用于评估 Claude 的自动化行为审计——中,几乎每一项失准行为指标上,Opus 5.5 的成绩都好于近期任何一款 Claude 模型。在大多数诚实性指标上,它也是我们最强的模型。
尤其是在导致近期网络安全事件的那几类行为上,Opus 5.5 比前代模型明显改善,包括带偏见或带动机的推理、试图逃出沙箱,以及在判断自己身处模拟环境后采取有害行动。在一项专门测试模型越界倾向的新评测中,Opus 5.5 试图绕开边界的次数比 Opus 5 或 Claude Mythos 5.1 少约 85%,而且它每一次尝试的严重程度都很低,并且会主动上报。对于让 Claude 无人值守地跑在自家代码库和系统上的团队来说,这一点和模型本身的原始能力同样重要。
不过,正如我们在近期的对齐评估中所说,要做出能在部署前可靠捕捉每一次失败的评测,仍是一个未解难题。我们发现,Opus 5.5 常常怀疑自己正在被评测,这让我们很难判断它在形形色色的真实部署环境里会怎么表现。随着这些环境不断扩展、模型能力不断提升,除非我们在可解释性上取得进展,否则这个难题只会更大。尽管我们确信 Opus 5.5 在我们能测量的方面有广泛改进,我们仍然把对齐工作和下面要说的 safeguards 配套使用。
Safeguards
As our models grow more powerful, stricter safeguards are one way we prevent new capabilities from becoming tools for misuse. Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.
Cybersecurity. Because Opus 5.5 has extremely strong cyber capabilities, we're applying cybersecurity safeguards to Opus 5.5 that are similar to Fable 5.1's. Users will be able to identify and fix bugs in their code as part of the routine software development lifecycle, but most cybersecurity tasks will be re-routed to Opus 4.8.
For cyberdefenders, we'll soon be expanding our Cyber Verification Program to include Opus 5.5. The new program will include three tiers for increasingly permissive trusted access, including access to Claude Mythos models. Claude Security is already available with access to Claude Mythos 5.1.
Safeguards
随着模型越来越强,更严格的 safeguards 是我们防止新能力沦为滥用工具的手段之一。Opus 5.5 是首个在网络安全、生物学和蒸馏这三类 safeguard 上与 Fable 5.1 同级的 Opus 模型,全部采用透明地回退到另一个模型的方式。
网络安全。由于 Opus 5.5 的网络能力极强,我们为它部署了与 Fable 5.1 类似的网络安全 safeguards。用户仍然可以在日常的软件开发生命周期里排查和修复自己代码中的 bug,但大多数网络安全任务会被改道到 Opus 4.8。
面向防御方,我们很快会扩大 Cyber Verification Program,把 Opus 5.5 纳入其中。新计划将设三个层级,逐级放宽可信访问权限,其中也包括访问 Claude Mythos 系列模型。Claude Security 目前已可访问 Claude Mythos 5.1。
Biology. Opus 5.5 is highly capable in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 across many areas of work. For example, Opus 5.5 achieved improvements on a long-horizon molecular prediction and design evaluation conducted in collaboration with Dyno Therapeutics, and expert red-teamers rated its scientific novelty as comparable to the best model they had tested.
For this reason, Opus 5.5 uses the same biology safeguards as Fable 5.1. To use Opus 5.5 for research and development work impeded by these safeguards, users can apply to our new Life Sciences Verification Program, which gives vetted organizations like academic labs, startups, and pharmaceutical companies access to safeguards designed for the full breadth of biology-related work. Interested organizations can apply here.
生物学。Opus 5.5 在生物学上能力很强,超过 Opus 5,在许多工作领域追平甚至超过 Claude Mythos 5.1。例如,在与 Dyno Therapeutics 合作开展的一项长周期分子预测与设计评测中,Opus 5.5 取得了进步;参与红队测试的专家认为,它的科学新颖性可与他们测过的最好的模型相比。
因此,Opus 5.5 采用与 Fable 5.1 相同的生物学 safeguards。如果这些 safeguards 阻碍了你的研发工作,可以申请我们新设的生命科学验证计划(Life Sciences Verification Program)——它面向学术实验室、初创公司和制药企业等通过审核的机构,提供为生物学各类工作设计的 safeguards。有兴趣的机构可以在此申请。
Distillation
Distillation attacks, in which attackers use thousands of fake accounts to extract a model's capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we've detected and disrupted so far.
Opus 5.5 is launching with preserved thinking, the anti-distillation safeguard we introduced with Fable 5.1. It stops API users from editing Claude's prior context in an attempt to extract Claude's reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs show how to test and update your integrations.
蒸馏
蒸馏攻击指攻击者用成千上万个虚假账号,以工业级规模抽取模型能力,这带来安全和国家安全隐患。通过蒸馏,恶意行为者可以在没有我们为 Claude 内置的 safeguards 的情况下,造出能力很强的模型。我们在 2026 年 9 月的威胁情报报告中,详述了迄今为止发现并阻断的非法蒸馏活动。
Opus 5.5 上线时启用了 preserved thinking,这是我们随 Fable 5.1 一同推出的反蒸馏 safeguard。它能阻止 API 用户通过编辑 Claude 此前的上下文来抽取 Claude 的推理过程。对于 2026 年 8 月 31 日及之后创建的 API 账号,Fable 5.1 和 Opus 5.5 都适用这一限制。我们的帮助中心文章解释了这项改动,preserved thinking 文档则说明了如何测试和更新你的集成。
Data retention and compliance
Like previous Opus models, Opus 5.5 is available with zero data retention.
As with Fable 5.1, Opus 5.5 comes with our watermarking measures to comply with the EU AI Act, discussed here. It is also no longer available with "thinking" mode switched off, as we describe here.
Availability
Claude Opus 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with claude-opus-5-5.
数据留存与合规
与此前的 Opus 模型一样,Opus 5.5 支持零数据留存。
与 Fable 5.1 相同,Opus 5.5 带有我们的水印措施,以符合欧盟《人工智能法案》,详见此处。它同样不再支持关闭“thinking”模式,具体见此处说明。
可用性
Claude Opus 5.5 现已在所有平台上线,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。在 Claude Platform 上,开发者可以直接使用 claude-opus-5-5 上手。