九月 01, 19:29
Anthropic 推出 Claude Fable 5.1,Fable 5 基準測試成績增加逾一倍
Anthropic Ships Claude Fable 5.1, More Than Doubling Its Predecessor on Key Benchmark
Decrypt

Anthropic 週二發布 Claude Fable 5.1 與 Claude Mythos 5.1。這些版本是 Fable 5 於 6 月 9 日推出後,Mythos 系列模型的首次更新。Fable 5.1 在 Terminal-Bench-Science 0.1 中取得 52.6% 的成績。Fable 5 在該基準測試中的成績為 24.7%。Opus 5 的成績為 29.0%。Terminal-Bench-Science 0.1 用於衡量 AI 代理能否在命令列環境中執行科學研究任務。Fable 5.1 在 Terminal-Bench 4.0 中取得 55.8% 的成績。Fable 5 在該測試中的成績為 42.0%。Opus 5 的成績為 52.3%。在較寬鬆的網路安全過濾條件下,Mythos 5.1 在 Terminal-Bench 4.0 中取得 60.9% 的成績。Anthropic 表示,其安全防護機制攔截了部分任務,並將其重新路由至 Opus 4.8。Fable 5.1 在不使用外部工具的情況下,於 Humanity's Last Exam 中取得 60.9% 的成績。使用外部工具時,Fable 5.1 的成績為 65.0%。Anthropic 表示,該模型是其用於學術目的的最先進模型。Anthropic 稱,新模型是全球最先進的程式設計和知識工作模型。Anthropic 表示,Fable 5.1 與 Mythos 5.1 使用相同的底層模型,但採用不同的安全過濾器。任何擁有 Claude 帳戶的人都可以使用 Fable 5.1。Mythos 5.1 則僅限於透過 Anthropic 的 Cyber Verification Program 和 Life Sciences Verification Program 審核的網路安全及生命科學專業人士。這些計畫取代了限制 Mythos 5 存取權限的 Project Glasswing 存取管道。Fable 5.1 的使用費用從 Pro 方案以及標準 Team 或 Enterprise 席位的按量付費使用額度中扣除。Fable 5.1 不計入這些方案的每週限額。Max 方案以及高級 Team 或 Enterprise 席位包含 Fable 5.1 的使用額度,最高相當於每週使用量的 50%。Fable 5.1 的價格為輸入 token 每個 $10、輸出 token 每個 $50。Opus 5 的價格為輸入 token 每個 $5、輸出 token 每個 $25。Anthropic 表示,在 Fable 5.1 上使用低或中等推理強度,可以用更低成本達到或超過 Fable 5 先前的成績。Fable 5.1 可透過 Claude API、Amazon Bedrock、Google Cloud 和 Microsoft Foundry 使用。其模型 ID 為 claude-fable-5-1。
This content is an AI-generated summary/analysis for informational purposes only and does not constitute investment advice.