Balyasny: evaluating and governing frontier models at $38B scale
An interview with Balyasny Asset Management's chief AI officer on how a $38B multi-strategy firm puts frontier models into production. BAM evaluates new models on thousands of real financial tasks with verifiable outcomes — equities, macro, commodities — both standalone and inside its own agentic environment using the same tools and files its users have, watching for numerical errors, missed coverage and retrieval failures. On the relevant subset Claude Fable 5 scored 89.4% versus 86.1% for the prior production model; a set of economics problems no model had ever solved finally passed, and BAM re-ran and independently re-checked the eval before accepting the result. Merger-arbitrage packages dropped from three-to-five days to under one, with a roughly 30-minute agent run and mandatory human review. Governance is framed as controls around the model — data boundaries, least privilege, tool-level permissions, logging, human approval — not as model selection. Note this is vendor-published customer material.
Balyasny Asset Management (BAM) is a global, multi-strategy investment firm that manages roughly $38 billion in assets and supports a team of roughly 2,000 investment professionals and staff. Charlie Flanagan, Chief AI Officer, spoke with Anthropic about how the firm evaluates new models on thousands of real financial tasks, why it built its own platform for running agents, and what changed with the launch of Claude Fable 5.
Balyasny Asset Management(BAM)是一家全球多策略投资公司,管理约 380 亿美元资产,拥有约 2,000 名投资专业人士和员工。首席 AI 官 Charlie Flanagan 与 Anthropic 探讨了该公司如何在数千个真实金融任务上评估新模型、为何自建 Agent 运行平台,以及 Claude Fable 5 的发布带来了哪些变化。
2026 is the year we moved from AI systems that do search to AI systems that do work.
The key change has been not just the models, but the harnesses around them, such as Claude Code. They allow the AI solutions we build to take on vastly more complex and longer-running tasks. The ability to give an AI an outcome rather than a prompt, and have it keep working until complete, has been a gamechanger.
A practical example is merger-arbitrage analysis. When a deal is announced, agents now build the initial deal-analysis package. They estimate how likely the deal is to close and how long it will take, extract the key economic and legal terms, identify conditions and milestones, and flag the areas that need investor judgment.
A year ago, those steps were fragmented across manual research and separate tools; we did not have an agent that could reliably sustain the full multi-step workflow to a usable conclusion. That work used to take three to five days. Now it takes less than one. The agent runs for approximately 30 minutes, with human review before any material output is relied on.
2026 年,我们从搜索型 AI 系统转向了工作型 AI 系统。
关键变化不仅在于模型本身,还在于围绕它们的执行框架,比如 Claude Code。这些框架让我们构建的 AI 解决方案能够承担远比以往复杂、运行时间更长的任务。能够给 AI 一个目标而非一段提示,并让它持续工作直到完成,这彻底改变了局面。
一个实际的例子是并购套利分析。当一笔交易宣布时,Agent 现在会构建初步的交易分析包。它们估算交易完成的可能性及所需时间,提取关键的经济和法律条款,识别条件和里程碑,并标出需要投资者判断的领域。
一年前,这些步骤还分散在人工研究和各种独立工具中;我们没有一个 Agent 能够可靠地完成完整的多步骤工作流并得出可用的结论。这项工作过去需要三到五天。现在不到一天。Agent 运行大约 30 分钟,任何重要输出在依赖之前都有人工审核。
We use Anthropic's frontier models, but most of the infrastructure is built in-house at BAM, including the execution harness, data access, and review controls.
我们使用 Anthropic 的前沿模型,但大部分基础设施是在 BAM 内部构建的,包括执行框架、数据访问和审核控制。
One thing we did years ago which has served us extremely well was invest in robust evaluation systems. We test new models on thousands of real-world financial tasks with verifiable outcomes, across equities, macro, and commodities, rather than relying on general benchmarks or isolated demonstrations. It has allowed us to make data-driven decisions around model choice and routing, and is something I think all enterprises should invest in.
We test both the model on its own and how it performs inside our agentic environment, with the same tools, files, and requirements our users have. Can it plan the work, choose and use the right tools, find and analyze evidence, recover from errors, check its intermediate results, and produce a grounded deliverable? We also look for specific failure modes, like numerical errors, missed coverage, unsupported conclusions, and retrieval problems.
多年前我们做了一件事,至今受益匪浅,那就是投资建设强大的评估体系。我们在数千个具有可验证结果的真实金融任务上测试新模型,覆盖股票、宏观和大宗商品,而不是依赖通用基准或孤立的演示。这让我们能够基于数据做出模型选择和路由决策,我认为所有企业都应该在这方面投入。
我们既测试模型本身,也测试它在我们的 Agent 环境中的表现,使用与用户相同的工具、文件和要求。它能规划工作、选择并使用正确的工具、查找和分析证据、从错误中恢复、检查中间结果,并产出有依据的交付物吗?我们还会寻找特定的故障模式,比如数值错误、覆盖遗漏、无根据的结论和检索问题。
On the relevant subset, Fable achieved 89.4% versus 86.1% for the prior production model, across thousands of tasks. Where it stood out most was complex planning, analysis, and agentic execution.
The surprising result was a set of economics problems we have tested that we have never had a model complete successfully, until Fable. We initially treated the result as a potential evaluation issue because it represented a material step change versus every model we had tested. We reran the evaluation, independently checked the task and scoring logic, and reviewed the result with Anthropic before concluding that the improvement was real. It was a wow moment.
Today, our investment teams use Fable as their go-to frontier model for systematic and coding work. We give them guidance on when to use Fable versus other models, based on efficiency and cost.
在相关子集上,Fable 在数千个任务中取得了 89.4% 的成绩,而之前的生产模型为 86.1%。它最突出的地方是复杂规划、分析和 Agent 执行。
令人惊讶的结果是一组我们测试过的经济学问题,此前从未有模型能成功完成,直到 Fable 出现。我们起初认为这可能是评估问题,因为它相对于我们测试过的所有模型都代表着实质性的飞跃。我们重新运行了评估,独立检查了任务和评分逻辑,并与 Anthropic 一起审查了结果,最终确认这一进步是真实的。那一刻令人惊叹。
如今,我们的投资团队将 Fable 作为系统化和编码工作的首选前沿模型。我们根据效率和成本,指导他们何时使用 Fable 以及其他模型。
We treat safety as a product and operating-model question, not as a one-time model-selection exercise. The relevant questions are not only what the model can do, but what data it can access, what tools it can use, what actions it can take, what must remain human-approved, and how we will know when something has gone wrong.
That means putting controls around the model rather than assuming the model itself is the control. We use approved data boundaries, least-privilege access, tool-level permissions, logging and traceability, human review for material outputs, and clear escalation paths for edge cases. We also test adversarial and failure scenarios before broadening access.
Those controls were a day-one priority, and security did not fundamentally change with Fable. A more capable model does not receive broader authority simply because it can reason or plan more effectively. Models can use only the tools and data sources approved for that user and task, and they cannot grant themselves more access. Investment judgment and accountability remain with people.
我们把安全视为产品和运营模式问题,而不是一次性的模型选择练习。关键问题不仅在于模型能做什么,还在于它能访问哪些数据、使用哪些工具、采取哪些行动、哪些必须经过人工批准,以及我们如何知道出了问题。
这意味着要在模型周围设置控制,而不是假设模型本身就是控制。我们采用批准的数据边界、最小权限访问、工具级权限、日志和可追溯性、重要输出的人工审核,以及边缘情况的明确升级路径。在扩大访问权限之前,我们还会测试对抗性和故障场景。
这些控制从第一天起就是优先事项,安全性并没有因为 Fable 而发生根本改变。一个能力更强的模型不会仅仅因为它能更有效地推理或规划就获得更广泛的权限。模型只能使用为该用户和任务批准的工具和数据源,并且不能为自己授予更多访问权限。投资判断和问责仍然由人负责。
Looking ahead, the direction of travel is toward more capable agents that can take longer-running, multi-step actions. That makes governance more important, not less. For us, that has meant building BAMAgent, our internal platform for securely deploying agents into approved enterprise workflows. It gives agents the tools and systems they need, but only those tools and systems. We have been building it for six months now and it supports thousands of autonomous agents working 24/7.
BAMAgent is the next step beyond our chat platform. Chat helps people take in and synthesize information. BAMAgent does the work: multi-step research and analysis that can run for hours or days, with agents working in parallel, and it ends in something a person can review. It can build and maintain a company research package, prepare for an earnings or macro event, or turn new evidence into financial scenarios. The agent plans the work, uses approved internal systems, runs the analysis, checks its intermediate outputs, and returns a research artifact, model, or decision-support package.
展望未来,趋势是朝着能力更强、能执行更长时间、多步骤行动的 Agent 发展。这让治理变得更加重要,而不是更不重要。对我们来说,这意味着构建 BAMAgent,这是我们内部用于将 Agent 安全部署到已批准的企业工作流中的平台。它为 Agent 提供所需的工具和系统,但仅限于这些工具和系统。我们已经构建了六个月,它支持数千个自主 Agent 全天候工作。
BAMAgent 是我们聊天平台之上的下一步。聊天帮助人们吸收和综合信息。BAMAgent 则完成工作:可以运行数小时或数天的多步骤研究和分析,Agent 并行工作,最终产出可供人审核的内容。它可以构建和维护公司研究包、为财报或宏观事件做准备,或将新证据转化为财务情景。Agent 规划工作、使用已批准的内部系统、运行分析、检查中间输出,并返回研究工件、模型或决策支持包。
Fable is our preferred model for the planning and analysis stages. A mistake there flows through every deliverable that follows, so we want the strongest available model deciding how to break down a problem, which evidence matters, and how to reconcile conflicting signals. That is what lets the agent work like a capable coworker.
Every enterprise should be developing a strategy to move toward a hosted-agent model that allows enterprise management and enforcement while maximizing the utility of agents for users.
Fable 是我们在规划和分析阶段的首选模型。这里的错误会影响后续每一个交付物,所以我们希望由可用的最强模型来决定如何分解问题、哪些证据重要,以及如何调和相互矛盾的信号。这正是让 Agent 像一位能干的同事那样工作的原因。
每家企业都应该制定战略,朝着托管 Agent 模式迈进,既实现企业管理和执行,又最大化 Agent 对用户的效用。
The reaction has been incredibly positive. Ultimately, people care about what this technology can unlock in their day-to-day work.
In one example, a BAMAgent ran a tax-loss harvesting analysis. It explored 90,000 database tables, found the relevant mutual fund holdings data, and built its own weighting system. After a review by our team, the result was more comprehensive than what a traditional approach would have produced.
Separately, our Chief Economist has configured an agent workflow that reduces a recurring central-bank analysis from roughly two days to approximately 30 minutes, with the economist retaining review and judgment.
反响非常积极。归根结底,人们关心的是这项技术能在日常工作中释放什么。
举个例子,一个 BAMAgent 运行了税务亏损收割分析。它探索了 9 万个数据库表,找到了相关的共同基金持仓数据,并构建了自己的加权系统。经过我们团队的审核,结果比传统方法所能产生的更为全面。
另外,我们的首席经济学家配置了一个 Agent 工作流,将一项经常性的央行分析从大约两天缩短到约 30 分钟,同时经济学家保留审核和判断。
Fable contributes the reasoning, synthesis, and multi-step problem-solving. BAM's harness provides the workflow design, approved data and tool access, retrieval context, permissions, monitoring, and human-review controls. Both are necessary for a production-quality result.
Fable 贡献推理、综合和多步骤问题解决能力。BAM 的执行框架提供工作流设计、经批准的数据和工具访问、检索上下文、权限、监控和人工审核控制。两者对于生产级质量的结果都必不可少。
This is the year we go from people having tools to having teammates. Much like a teammate, agents will become more useful over time as you work with them, complete more complex tasks, and start to do work proactively to help.
We already have some teams running over 300 agents doing analysis over new data and information constantly. It helps the teams both be faster to insights and not miss anything.
It also changes the question we ask. We used to build expert systems and teach people to automate the processes they already had. Now we ask whether there is a better way to reach the outcome.
今年我们将从人们拥有工具,转变为拥有队友。就像队友一样,随着你与 Agent 合作、完成更复杂的任务,并开始主动帮忙,它们会变得越来越有用。
我们已经有团队运行着 300 多个 Agent,不断分析新的数据和信息。这帮助团队更快获得洞察,同时不遗漏任何东西。
这也改变了我们提出的问题。我们过去构建专家系统,教人们自动化已有的流程。现在我们问的是,有没有更好的方式达成结果。
The limits are really just our own imagination. The tools, data, and models are now at a point where they can do real work for hours on end; it's up to us to continue to reimagine what is possible. It is going to be an incredibly exciting next 12 months.
Get started with Claude Fable.
限制其实只在于我们自己的想象力。工具、数据和模型如今已经达到可以连续数小时做真正工作的程度;接下来要靠我们继续重新想象什么是可能的。未来 12 个月将令人无比兴奋。
开始使用 Claude Fable。