Glean 拾遗
Daily /2026-07-20 / Interview with Yang Zhilin: AGI, Long Context, and China's AI Race

Interview with Yang Zhilin: AGI, Long Context, and China's AI Race

Source x.com Glean’d 2026-07-20 06:00 Read 60 min
AI summary

In early 2024, Yang Zhilin, founder of Moonshot AI (Kimi), gave an in-depth interview covering his AGI conviction, the strategic choice of long context as the first step, views on open-source vs closed-source, and the organizational form needed for AI startups. He argues that true AGI companies must combine science, engineering, and business, insist on B2C, and synchronize user scaling with model scaling. The interview includes real-time assessments of Sora, catching up with GPT-4, and the gap between Chinese and Silicon Valley AI companies, revealing the balance between technical idealism and pragmatism.

Original · 60 min
x.com ↗
§ 1

A Conversation with Yang Zhilin of Kimi: Advancing Toward the Endless, Unknown Snow Mountains

与Kimi杨植麟对话:朝向无尽未知的雪山

§ 2

This was my first interview with Yang Zhilin, conducted in early 2024 and published on March 1, 2024—exactly the first anniversary of Kimi's founding. At the time, Kimi had only 80 people, working out of their first, somewhat run-down office. There was no logo at the entrance. Only a white piano standing guard by the door.

This article generated quite a stir in China's tech community at the time.

Back then, I was still a print journalist, so this interview exists only as text and an audio podcast.

As we can see, many of the views expressed in this article have since been borne out.

Just rereading these words from two years ago, you really can't help but marvel at how dramatically the world has changed!

(This translation was generated by Kimi K3.)

这是我与杨植麟的第一次采访,于2024年初进行,并于2024年3月1日发布——恰逢Kimi成立一周年。当时,Kimi只有80人,在第一个略显破旧的办公室里工作。入口处没有标志,只有一架白色钢琴守护在门旁。

这篇文章当时在中国科技界引起了不小的轰动。

那时我还是一名纸媒记者,因此这次采访仅以文字和音频播客形式存在。

正如我们所见,本文中的许多观点后来都得到了印证。

重读两年前的这些文字,你不得不感叹世界的变化有多大!

(本翻译由Kimi K3生成。)

§ 3

Yang Zhilin: “If everyone thinks you are normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.”

杨植麟:“如果所有人都认为你正常——如果你的梦想是任何人都能拥有的——那它就不会增加人类梦想的总和。”

§ 4

By Zhang Xiaojun

Just one year ago, AI scientist Yang Zhilin did a precise calculation in Silicon Valley. He realized that if he decided to launch a foundation-model startup aimed at AGI, he would need to raise more than $100 million within the next few months.

Yet that was merely a ticket to the game. A year later, that figure had multiplied thirteenfold.

For foundation-model companies, competition is less a scientific contest than, first and foremost, a brutal contest of money. With investors holding their purse strings tight, you have to stay ahead of your rivals in raising more money, buying more GPUs, and grabbing more talent.

“It requires a concentration of talent and a concentration of capital,” says Yang Zhilin, founder and CEO of Moonshot AI, the foundation-model company established on March 1, 2023.

Over the past year, Chinese foundation-model companies have seemed to live on a tense, constricted edge of survival. On the surface, each of them holds hefty sums of cash. But on one hand, they must immediately pour freshly raised money into extremely costly research to chase OpenAI—first catching up to GPT-3.5, and before GPT-4 is even reached, along comes Sora. On the other hand, they must race nonstop to find viable real-world use cases, to validate for themselves that they are companies, not research institutes that only devour capital. And that is not enough: for every one of these ventures, whether the exit is an IPO or an acquisition, the way out remains anything but clear.

Among the founders of China’s foundation-model companies, Yang Zhilin is the youngest, born in 1992. The industry describes him as a staunch AGI believer and a founder with rare technical charisma. Much of his academic and professional record is tied to general-purpose AI, and his papers have been cited more than 22,000 times.

In mid-2023, China’s tech community turned abruptly from euphoria to chill on foundation models, and accelerating real-world deployment became the pragmatic mainstream melody. This inevitably left foundation-model CEOs torn violently between ideal and reality. In a Chinese AI ecosystem where everyone chants PMF (product/market fit) and everyone chants commercialization, this founder—an AI researcher by training—is in no particular hurry.

With 80 people, Moonshot AI has the smallest headcount among China’s leading foundation-model companies. Unlike his rivals, Yang did not opt for the safer B2B business or seek deployment in verticals such as healthcare or gaming. He built one—and only one—consumer product: the AI assistant Kimi, which accepts inputs of up to 200,000 Chinese characters. Kimi is also Yang Zhilin’s English name.

Yang prefers to see his company as a system that combines science, engineering, and business. You might picture it this way: above the human world, he is erecting an AI laboratory bench—with one hand he runs experiments, and with the other he brings cutting-edge technology down into the real world, discovering applications through interaction with people and delivering those applications into consumers’ hands. Ideally, the former burns through billions and tens of billions of dollars of capital; the latter earns that money back hundreds or thousands of times over. However you hear it, it sounds as thrilling—and as perilous—as walking a tightrope.

“AI is not about what PMF I can find in the next year or two; it’s about how to change the world over the next ten to twenty years,” he says.

Such abstract, idealistic thinking makes one sweat nervously on his behalf: can a young AI scientist carve out room to survive in a realist China?

In February 2024, Moonshot AI closed a large funding round against the market tide. It is understood that the company raised a Series B of more than $1 billion at a $1.5 billion pre-money valuation, led by Alibaba with follow-on participation from Monolith Management, Xiaohongshu, and others. Upon completion of the deal, Moonshot AI’s post-money valuation stood at roughly $2.5 billion—making it, at this stage, the highest-valued unicorn in China’s foundation-model race. (The company declined to respond to or comment on the matter.)

In the midst of this third funding round, we sat down with Yang Zhilin to talk about his first year of entrepreneurship—a cross-section, in miniature, of a year in which Chinese foundation-model companies raced ahead from the starting line.

His company did not set up in Sohu Network Plaza in Beijing, the hub where foundation-model companies cluster. For a company with total funding of about RMB 9 billion, this office in the Liangzi Xinzuo building looks crude and run-down. There is not even a company logo at the entrance—only a white piano standing guard by the door.

The meeting room sits in a corner; with its small windows it is dark inside, and the heater hums as it blows warm air against the winter cold. In the dim light, Yang describes how the past year has felt to him: “It’s a bit like driving down a road with a range of snow mountains stretching out ahead. You don’t know what’s inside them. You just keep walking forward, one step at a time.”

Below is the full interview with Yang Zhilin. (For readability, the author has made some textual edits.)

作者:张晓军

就在一年前,AI科学家杨植麟在硅谷做了一个精确计算。他意识到,如果决定创办一家以AGI为目标的基础模型初创公司,他需要在接下来几个月内筹集超过1亿美元。

然而,这只是入场券。一年后,这个数字翻了13倍。

对于基础模型公司而言,竞争与其说是科学竞赛,不如说首先是残酷的金钱竞赛。投资者收紧钱袋,你必须在筹集更多资金、购买更多GPU、抢夺更多人才方面领先对手。

“这需要人才和资本的集中,”月之暗面(Moonshot AI)创始人兼CEO杨植麟说。该公司于2023年3月1日成立。

过去一年,中国基础模型公司似乎生活在紧张而局促的生存边缘。表面上看,每家公司都持有巨额现金。但一方面,他们必须立即将新融资投入成本高昂的研究,以追赶OpenAI——先追上GPT-3.5,还没达到GPT-4,Sora又来了。另一方面,他们必须不断寻找可行的现实用例,以验证自己是公司,而非只消耗资本的研究机构。这还不够:对每家公司来说,无论退出是IPO还是收购,出路依然不明朗。

在中国基础模型公司创始人中,杨植麟最年轻,1992年出生。业界将他描述为坚定的AGI信徒和拥有罕见技术魅力的创始人。他的大部分学术和职业经历都与通用AI相关,论文被引用超过22,000次。

2023年中,中国科技界对基础模型的态度从狂热骤然转为冷静,加速实际部署成为务实主流。这不可避免地让基础模型CEO们在理想与现实间剧烈撕裂。在一个人人高喊PMF(产品市场契合)和商业化的中国AI生态中,这位受过AI研究员训练的创始人并不着急。

月之暗面仅有80人,是中国领先基础模型公司中人数最少的。与对手不同,杨植麟没有选择更安全的B2B业务或寻求在医疗、游戏等垂直领域落地。他构建了唯一一个消费品:AI助手Kimi,可接受多达20万中文字符的输入。Kimi也是杨植麟的英文名。

杨植麟更愿意将公司视为一个结合科学、工程和商业的系统。可以这样想象:在人类世界之上,他正在搭建一个AI实验台——一只手做实验,另一只手将前沿技术带入现实世界,通过与人的互动发现应用,并将这些应用交付到消费者手中。理想情况下,前者消耗数十亿甚至数百亿美元资金;后者则将这些钱成百上千倍地赚回来。无论如何,这听起来像走钢丝一样刺激——也危险。

“AI不是关于我未来一两年能找到什么PMF,而是关于未来十到二十年如何改变世界,”他说。

这种抽象而理想主义的思考让人为他捏把汗:一个年轻的AI科学家能否在现实主义的中国找到生存空间?

2024年2月,月之暗面逆市完成一轮大额融资。据悉,公司以15亿美元投前估值完成超过10亿美元的B轮融资,由阿里巴巴领投,思源资本、小红书等跟投。交易完成后,月之暗面投后估值约25亿美元——使其成为当前中国基础模型竞赛中估值最高的独角兽。(公司拒绝回应或评论此事。)

在第三轮融资进行中,我们与杨植麟坐下来,聊了聊他创业的第一年——这也是中国基础模型公司从起点起跑的一年缩影。

他的公司没有设在北京基础模型公司聚集的搜狐网络大厦。对于一家总融资约90亿人民币的公司来说,这个在亮子新座大楼里的办公室显得简陋破旧。入口处甚至没有公司标志——只有一架白色钢琴守护在门旁。

会议室在角落,小窗户使得室内昏暗,暖气嗡嗡作响,吹着热风抵御冬寒。在昏暗的光线下,杨植麟描述过去一年的感受:“有点像开车走在一条路上,前方是一片雪山。你不知道里面是什么。你只能一步步往前走。”

以下是杨植麟的完整采访。(为便于阅读,作者做了部分文字整理。)

§ 5

§ 6

Part 1

Standing at the Beginning

“You Have to Ride the Wave”

Zhang Xiaojun: How have you been lately?

Yang Zhilin: Busy—there’s a lot going on. But I’m still excited. We’re standing at the very beginning of an industry, and there is enormous room for imagination.

Zhang Xiaojun: When I came in just now, I saw a pure white piano at your company’s entrance.

Yang Zhilin: There’s a Pink Floyd album sitting on it, too. I have no idea who put them there—I suddenly noticed them a couple of days ago and haven’t had a chance to ask. (Pink Floyd is the British rock band that released the album The Dark Side of the Moon.)

第一部分

站在起点

“你必须乘上浪潮”

张晓军:你最近怎么样?

杨植麟:忙碌——有很多事情。但我仍然兴奋。我们正处在一个行业的起点,想象空间巨大。

张晓军:刚才进来时,我看到公司门口有一架纯白色钢琴。

杨植麟:上面还有一张平克·弗洛伊德的专辑。我不知道是谁放的——几天前突然注意到,还没机会问。(平克·弗洛伊德是发行专辑《月之暗面》的英国摇滚乐队。)

§ 7

Zhang Xiaojun: On the day ChatGPT was released in November 2022, what were you doing?

Yang Zhilin: I was already preparing for this—recruiting people, building a team, exchanging new ideas. Seeing ChatGPT was thrilling. Three to five years earlier—even in 2021—it would have been inconceivable. That kind of higher-order reasoning had been very hard to achieve.

I sensed that many variables were about to shift in the market: capital on one side, talent on the other—the core factors of production for AI. If those variables fell into place, it would become possible to build a proper company to do this—an organization built for AGI could go from 0 to 1. That was a major epiphany. An independent company made more sense, but it wasn’t something you could do the moment you wanted to; ChatGPT jolted the variables and brought the factors of production together. You have to ride the wave.

张晓军:2022年11月ChatGPT发布那天,你在做什么?

杨植麟:我已经在为此做准备——招人、组建团队、交流新想法。看到ChatGPT很激动。三五年前,甚至2021年,这都是无法想象的。那种高阶推理能力一直很难实现。

我感觉到市场中的许多变量即将变化:一方面是资本,另一方面是人才——AI的核心生产要素。如果这些变量到位,就有可能建立一家合适的公司来做这件事——一个为AGI构建的组织可以从0到1。这是一个重大顿悟。独立公司更有意义,但并不是你想做就能做;ChatGPT冲击了变量,使生产要素聚集。你必须乘上浪潮。

§ 8

Zhang Xiaojun: After you decided to found an AGI company, what preparations did you make? How did you assemble the two factors of production—capital and talent?

Yang Zhilin: It was a winding process. ChatGPT took time to diffuse. Some people learned of it early, some late; some doubted at first, then were shocked, then became believers. Finding people and finding money were tightly bound to timing.

We began focusing on our first funding round in February 2023. Had we delayed to April, we basically would have had no chance. But doing it in December 2022 or January 2023 wouldn’t have worked either—the pandemic was still on, and people hadn’t processed it yet. So the real window was just one month.

One night in the United States, I did a precise calculation. When I finished, I concluded we needed to raise at least $100 million within a few months. Many in the market hadn’t started fundraising yet, and many didn’t believe you could necessarily raise that much. But it turned out to be possible—even more than that.

The talent market started moving, too. Inspired by ChatGPT, many people had this realization in March or April 2023: this is the only thing worth doing in the next decade. You have to reach out to the right people at the right time. A year or two earlier, talent would not have clustered to this degree. Back then, more people were doing traditional AI or AI-adjacent businesses—none of it was general-purpose AI.

Zhang Xiaojun: To sum up: February was the window for fundraising, and March and April were the window for hiring?

Yang Zhilin: More or less.

Zhang Xiaojun: That night in the U.S.—where were you when you did this math? How exactly did you calculate it?

Yang Zhilin: From late 2022 into early 2023, I spent a month or two in the U.S., talking to people. I did it where I was living. You work out how many FLOPs you need, the training cost, inference, and the user numbers.

张晓军:决定创办AGI公司后,你做了哪些准备?如何汇聚资本和人才这两个生产要素?

杨植麟:这是一个曲折的过程。ChatGPT的传播需要时间。有人早听说,有人晚听说;有人最初怀疑,后来震惊,再成为信徒。找人和找钱都与时机紧密相连。

我们在2023年2月开始聚焦首轮融资。如果拖到4月,基本上就没机会了。但2022年12月或2023年1月也不行——疫情还在,大家没反应过来。所以真正的窗口只有一个月。

在美国的一个晚上,我做了精确计算。算完后,我得出结论:我们需要在几个月内筹集至少1亿美元。市场上很多人还没开始融资,很多人也不相信能筹到那么多。但事实证明是可能的——甚至更多。

人才市场也开始流动。受ChatGPT启发,许多人在2023年3月或4月意识到:这是未来十年唯一值得做的事。你必须在合适的时间联系到合适的人。早一两年,人才不会如此聚集。那时更多人做传统AI或AI周边业务——都不是通用AI。

张晓军:总结一下:2月是融资窗口,3月和4月是招聘窗口?

杨植麟:差不多。

张晓军:美国那晚——你在哪里算的?具体怎么算的?

杨植麟:从2022年底到2023年初,我在美国待了一两个月,和人交流。我在住的地方算的。你需要算出多少FLOPs、训练成本、推理和用户数量。

§ 9

Zhang Xiaojun: At that moment, what was the prevailing mood in Silicon Valley?

Yang Zhilin: The product began picking up many early adopters, concentrated in the tech circle. We were in that circle ourselves, so we felt it more keenly. At the big Silicon Valley companies, people have to write performance reviews every six months, and many started writing them with ChatGPT. Some people whose writing was usually not that professional turned in reviews written with ChatGPT, and everyone sounded dead serious.

Undercurrents were stirring. Many people were thinking about their next job or about starting a company. Quite a few friends who talked with us later went off to found startups. And there was intense FOMO—fear of missing out. Nobody could sleep. Whether it was midnight, 1 a.m., or 2 a.m., if you reached out, people were always there. A bit anxious, a bit FOMO, and very excited.

张晓军:那时硅谷的主流情绪是什么?

杨植麟:产品开始吸引许多早期用户,集中在科技圈。我们就在这个圈子里,所以感受更强烈。在硅谷大公司,人们每半年要写绩效评估,很多人开始用ChatGPT写。一些平时写作不太专业的人交了ChatGPT写的评估,每个人听起来都极其严肃。

暗流涌动。很多人考虑下一份工作或创业。不少后来和我们聊过的朋友都去创办了初创公司。还有强烈的FOMO——害怕错过。没人能睡着。无论是午夜、凌晨1点还是2点,你发消息,人都在。有点焦虑,有点FOMO,又非常兴奋。

§ 10

Zhang Xiaojun: The night you calculated you needed to raise $100 million—how late did you stay up?

Yang Zhilin: It was fine—the calculation itself didn’t take long. But afterward, I couldn’t tell too many people. If I had, no one would have believed it could be done.

张晓军:算出需要1亿美元那晚,你熬夜到多晚?

杨植麟:还好——计算本身没花太久。但之后,我不能告诉太多人。如果说了,没人会相信能做成。

§ 11

Part 2

Technical Lineage

“Free Yourself from Endless Carving”

Zhang Xiaojun: When the venture capital world talks about you, they say, “The founder is brilliant, has technical charisma, and the team is full of technical stars.” So before we discuss your foundation-model venture, I’d like to start with your academic background. You studied computer science at Tsinghua as an undergraduate and earned your PhD at Carnegie Mellon’s School of Computer Science. Has AI always been your focus?

Yang Zhilin: I was born in 1992 and started my undergraduate degree in 2011. From my sophomore year to now—more than a decade—I’ve been in this field. At first I explored more divergently, looking around everywhere; I did some work related to graphs and to multimodality. In 2017, I converged on language models. At the time I felt language models were a relatively important problem; later I came to feel it was the only important problem.

Zhang Xiaojun: In 2017, how did the AI industry generally understand language models, and how did that understanding evolve?

Yang Zhilin: Back then it was a model used to rank speech-recognition results. (Laughs.) After a segment of speech was recognized, you’d get many candidate results, and you’d use the language model to see which one had the highest probability and output the most likely one. Its applications were very limited.

But you come to realize it’s a fundamental problem, because you are modeling the probabilities of the world. Language is limited, but it’s a projection of the world; in theory, if you make the token space—the space of all possible tokens—large enough, you can build a general world model. How everything in the world arises and develops can be assigned a probability. Every problem can be reduced to how to estimate probabilities.

第二部分

技术传承

“从无尽的雕琢中解脱”

张晓军:创投圈谈及你时,会说“创始人厉害,有技术魅力,团队全是技术明星”。所以在讨论你的基础模型创业之前,我想从你的学术背景开始。你在清华读本科,在卡内基梅隆大学计算机科学学院读博。AI一直是你的重点吗?

杨植麟:我1992年出生,2011年开始本科。从大二到现在——十多年——我都在这个领域。一开始我探索得更发散,到处看;做过一些与图和多模态相关的工作。2017年,我收敛到语言模型上。当时我觉得语言模型是一个相对重要的问题;后来我觉得它是唯一重要的问题。

张晓军:2017年,AI界对语言模型的理解是怎样的?这种理解后来如何演变?

杨植麟:当时它是一个用于对语音识别结果进行排名的模型。(笑。)一段语音被识别后,会有很多候选结果,用语言模型看哪个概率最高,输出最可能的那个。它的应用非常有限。

但你会意识到它是一个基本问题,因为你是在对世界的概率建模。语言是有限的,但它是世界的一个投影;理论上,如果你让token空间——所有可能token的空间——足够大,你就能构建一个通用的世界模型。世界上一切事物的发生和发展都可以赋予一个概率。每个问题都可以归结为如何估计概率。

§ 12

Zhang Xiaojun: Your academic mentors are very prominent: your PhD advisors were Ruslan Salakhutdinov, head of AI at Apple, and William W. Cohen, chief scientist of Google AI. Both straddle industry and academia.

Yang Zhilin: In previous years, industry and academia came together more, but the trend is now shifting: more valuable breakthroughs will happen in industry. That’s an inevitable law of development. It starts with exploratory research and gradually shifts into a more mature industrialization process. That doesn’t mean research is unnecessary during industrialization—only that pure research will struggle to produce valuable breakthroughs.

Zhang Xiaojun: What did you learn from these renowned mentors?

Yang Zhilin: I learned the most at Google, where I interned for a long time. I began working on Transformer-based language models in late 2018. My biggest learning was freeing myself from endless “carving”—the obsessive refinement of surface details. That was crucial.

You should look at what the big direction is, the big gradient. When ten roads lie before you, the average person worries about how to brake for a pedestrian ahead on this one road—short-term details. But which of the ten roads to take is what matters most.

This field previously had exactly that problem. For example, on a dataset of only one or two million tokens, you’d look at how to push perplexity lower, how to push loss lower, how to improve accuracy—and you’d fall into endless carving. People invented many bizarre architectures; these were carving tricks. After carving, you might do better on that kind of dataset, but you miss the essence of the problem.

The essence is analyzing what the field is missing. What is the first principle? Why can the scaling law serve as a first principle? You only need to find a structure that satisfies two conditions: first, it is sufficiently general; second, it is scalable. General means you can model all problems within this framework; scalable means that as long as you pour in enough compute, it keeps getting better.

This is the thinking I learned at Google: if something can be explained by something more fundamental, you shouldn’t over-carve at the upper layers. There’s an important line I strongly agree with: if you can solve a problem with scale, don’t solve it with a new algorithm. The greatest value of a new algorithm is in how it lets you scale better. When you free yourself from carving, you can see much more.

张晓军:你的学术导师非常杰出:博士导师是苹果AI负责人Ruslan Salakhutdinov和Google AI首席科学家William W. Cohen。他们都横跨产业和学术界。

杨植麟:过去几年,产业和学术界结合更紧密,但现在趋势在变:更有价值的突破将发生在产业界。这是发展的必然规律。从探索性研究开始,逐渐转向更成熟的工业化过程。这并不意味着研究在工业化过程中不必要——只是纯粹的研究将难以产生有价值的突破。

张晓军:你从这些著名导师那里学到了什么?

杨植麟:在Google学到的最多,我在那里实习了很长时间。2018年底开始研究基于Transformer的语言模型。最大的收获是摆脱无尽的“雕琢”——对表面细节的过度打磨。这很关键。

你应该看大方向、大梯度。当十条路在你面前时,普通人担心的是在这条路上如何为前面的行人刹车——短期细节。但走哪条路才是最重要的。

这个领域以前就有这个问题。比如,在只有一两百万token的数据集上,你关注如何降低困惑度、如何降低损失、如何提高准确率——然后陷入无尽雕琢。人们发明了许多奇特的架构;这些都是雕琢技巧。雕琢后,你在那种数据集上可能更好,但错过了问题的本质。

本质是分析这个领域缺少什么。第一性原理是什么?为什么规模法则可以作为第一性原理?你只需要找到一个满足两个条件的结构:第一,它足够通用;第二,它是可扩展的。通用意味着你可以在这一框架内对所有问题建模;可扩展意味着只要你投入足够的算力,它就会越来越好。

这就是我在Google学到的思维:如果某件事可以用更基本的东西解释,就不应该在更高层过度雕琢。有一句重要的话我强烈认同:如果你能用规模解决问题,就不要用新算法解决。新算法的最大价值在于如何让你更好地扩展。当你从雕琢中解脱出来,你能看到更多。

§ 13

Zhang Xiaojun: Was Google also a follower of the scaling law back then? How did it implement first-principles thinking?

Yang Zhilin: Many such ideas already existed there, but Google didn’t implement them especially well. It had this way of thinking, but it couldn’t organize itself into a true moonshot. It was more like: here are five people pursuing my first principles, and over there five people pursuing theirs. There was nothing top-down.

张晓军:Google当时也是规模法则的追随者吗?它是如何实践第一性原理思维的?

杨植麟:很多这类想法已经存在,但Google执行得不太好。它有这种思维方式,但无法将自己组织成一个真正的登月计划。更像是:这里有五个人追求我的第一性原理,那边有五个人追求他们的。没有自上而下的推动。

§ 14

Zhang Xiaojun: During your PhD, you published papers in collaboration with Turing Award winners Yann LeCun and Yoshua Bengio—and you were first author on those papers. How did those collaborations come about? What I mean is: they’re Turing Award laureates and they weren’t your advisors—what did you rely on to attract them?

Yang Zhilin: Academia is very open. As long as you have a good idea and a meaningful problem, it’s fine. What two brains—or n brains—produce is more than one brain alone. This applies when developing AGI, too. An important strategy in AI is called “ensembling”—using multiple different models or methods and combining their predictions for better performance. It’s essentially doing the same thing: when you have diverse viewpoints, you can spark many new things. Collaboration is hugely beneficial.

Zhang Xiaojun: Would you first have an idea and then ask them whether they were interested?

Yang Zhilin: That’s roughly how it went.

Zhang Xiaojun: Which is harder: winning over academic heavyweights in research, or winning over capital heavyweights in fundraising? What are the similarities?

Yang Zhilin: “Winning over” isn’t a good phrase—the essence behind it is cooperation. Cooperation means both sides win, because mutual benefit is the precondition for cooperation. So there’s really no difference: you need to offer others unique value.

Zhang Xiaojun: How do you earn their trust? What do you think your gift is?

Yang Zhilin: There’s no particular gift—just working hard.

张晓军:读博期间,你与图灵奖得主Yann LeCun和Yoshua Bengio合作发表了论文,并且你是第一作者。这些合作是如何发生的?我的意思是:他们是图灵奖得主,又不是你的导师——你靠什么吸引他们?

杨植麟:学术界非常开放。只要你有好想法和有价值的问题,就没问题。两个大脑——或n个大脑——产生的成果比一个大脑多。这在开发AGI时也适用。AI中一个重要的策略叫“集成”——使用多种不同模型或方法,结合它们的预测以获得更好性能。本质上是一样的:当你有不同观点时,可以激发许多新东西。合作大有裨益。

张晓军:你会先有想法,然后问他们是否感兴趣吗?

杨植麟:大致如此。

张晓军:哪个更难:在研究上赢得学术大咖,还是在融资上赢得资本大咖?有什么相似之处?

杨植麟:“赢得”不是个好词——它背后的本质是合作。合作意味着双方共赢,因为互利是合作的前提。所以没什么不同:你需要为他人提供独特价值。

张晓军:你如何赢得他们的信任?你认为自己的天赋是什么?

杨植麟:没什么特别天赋——只是努力工作。

§ 15

Part 3

The Old System No Longer Works

“AGI Needs a New Kind of Organization”

Zhang Xiaojun: You just said “more valuable breakthroughs will happen in industry”—does that include startups and the giants’ AI labs?

Yang Zhilin: Labs are history. Google Brain used to be the biggest AI lab in industry, but it was a research organization embedded inside a big company. That kind of organization can explore new ideas, but it’s very hard for it to produce a great system—it could produce the Transformer, but it couldn’t produce ChatGPT.

The way development now evolves is that you’re building an enormous system, which requires new algorithms, solid engineering, and even a lot of product and commercialization work. It’s like the early 2000s: you couldn’t research information retrieval in a lab; it had to live in the real world, as a huge system, a product with users—like Google. So research and education systems will shift their function toward primarily cultivating talent.

第三部分

旧系统不再有效

“AGI需要一种新型组织”

张晓军:你刚才说“更有价值的突破将发生在产业界”——这包括初创公司和巨头的AI实验室吗?

杨植麟:实验室已经是历史了。Google Brain曾经是产业界最大的AI实验室,但它是嵌入大公司的研究组织。这种组织可以探索新想法,但很难产生一个伟大的系统——它能产生Transformer,但无法产生ChatGPT。

现在的发展方式是,你正在构建一个巨大的系统,这需要新算法、扎实的工程,甚至大量产品和商业化工作。就像21世纪初:你不能在实验室里研究信息检索;它必须存在于现实世界中,作为一个巨大系统、一个有用户的产品——比如Google。因此,研究和教育系统将转向主要培养人才的功能。

§ 16

Zhang Xiaojun: How would you describe this new form of system? Is OpenAI its prototype?

Yang Zhilin: It’s the most mature organization of this kind today, and it’s still gradually evolving.

Zhang Xiaojun: So it can be understood as an organization established for humanity’s grand scientific goals?

Yang Zhilin: I want to emphasize: it is not pure science—it’s a combination of science, engineering, and business. It has to be a commercial organization, a company, not a research institute. But this company is built from zero to one, because AGI needs a new kind of organization. First, the mode of production differs from the internet era; second, it shifts from pure research to a combination of research, engineering, product, and business.

At its core, it should be a moonshot program, with a great deal of top-down planning—yet within that planning there is room for innovation, because not all the technology is predetermined. Bottom-up elements exist within a top-down framework. Such an organization didn’t exist before, but the organization must adapt to the technology, because technology determines the mode of production; if they don’t match, you can’t produce effectively. We believe it will very likely need to be redesigned from scratch.

张晓军:你如何描述这种新系统形式?OpenAI是它的原型吗?

杨植麟:它是目前这类组织中最成熟的,而且还在逐渐演变。

张晓军:所以可以理解为是为人类伟大科学目标而建立的组织?

杨植麟:我想强调:它不是纯粹的科学——它是科学、工程和商业的结合。它必须是商业组织、一家公司,而不是研究机构。但这家公司是从零到一构建的,因为AGI需要一种新型组织。首先,生产方式与互联网时代不同;其次,它从纯研究转向研究、工程、产品和商业的结合。

其核心应该是一个登月计划,有大量自上而下的规划——但在规划中有创新空间,因为并非所有技术都是预设的。在自上而下的框架内存在自下而上的元素。这种组织以前不存在,但组织必须适应技术,因为技术决定生产方式;如果不匹配,就无法有效生产。我们相信它很可能需要从头重新设计。

§ 17

Zhang Xiaojun: You want to build “China’s OpenAI”—can we put it that way?

Yang Zhilin: Not quite accurate. We don’t want to be China’s anything, and we don’t necessarily want to be OpenAI.

First, real AGI will definitely be global. There is no such thing—at least not long-term—as an AGI company confined to some regional market because of market-protection mechanisms. Globalization, AGI, and having a product with a very large user base: these three are ultimately necessary conditions.

Second, should it be OpenAI? If you look at 2017–2018, OpenAI had a terrible reputation. When people in our circle looked for jobs, they generally considered places like Google. Many people who talked with Ilya Sutskever, OpenAI’s chief scientist, came away thinking the man was crazy and far too full of himself—OpenAI was either madmen or scammers. But they committed very early, found the non-consensus, and found what is now the only first principle that works in AI: scaling through next-token prediction.

I believe there will be a company greater than OpenAI. A truly great company can combine technological idealism with a great product, co-creating with its users—AGI will ultimately be something produced by co-working with all of its users. So it’s not only about technology; it also requires pragmatism and real-world pursuits—ultimately, a perfect combination of the two.

Still, we should learn from OpenAI’s technological idealism. If everyone thinks you’re normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.

张晓军:你想打造“中国的OpenAI”——可以这么说吗?

杨植麟:不太准确。我们不想成为中国的任何东西,也不一定想成为OpenAI。

首先,真正的AGI必然是全球性的。不存在——至少长期不存在——由于市场保护机制而局限于某个区域市场的AGI公司。全球化、AGI、拥有大规模用户的产品:这三个最终是必要充分条件。

其次,应该是OpenAI吗?看看2017-2018年,OpenAI名声很差。我们圈子里的人找工作,通常考虑Google之类的地方。许多与OpenAI首席科学家Ilya Sutskever交谈过的人,都觉得这个人疯了,太自以为是——OpenAI要么是疯子,要么是骗子。但他们很早就承诺,找到了非共识,找到了现在AI中唯一有效的第一性原理:通过下一个token预测进行扩展。

我相信会有比OpenAI更伟大的公司。一个真正伟大的公司可以将技术理想主义与伟大产品结合,与用户共同创造——AGI最终将是与所有用户协同工作的产物。所以不仅是技术,还需要务实和对现实世界的追求——最终是两者的完美结合。

不过,我们应该学习OpenAI的技术理想主义。如果所有人都认为你正常——如果你的梦想是任何人都能拥有的——那它就不会增加人类梦想的总和。

§ 18

Part 4

The Moonshot’s First Step Is “Long Context”—What’s the Second?

“Two Big Milestones Are Coming Next”

Zhang Xiaojun: Back to the moment you decided to start the company—did you launch the first funding round immediately after returning to China?

Yang Zhilin: It began in the U.S. in February (last year), some of it remotely. In the end, domestic investors made up the majority.

Zhang Xiaojun: Did the first round raise $100 million?

Yang Zhilin: The first round wasn’t that much; later rounds exceeded that figure. We completed two rounds in 2023, totaling nearly RMB 2 billion.

This is now the third round. We haven’t formally announced the financing, so I can’t comment at this time.

Zhang Xiaojun: Some people say that since the second half of 2023, no one has been willing to invest in foundation-model companies anymore. Are they wrong?

Yang Zhilin: There still are. You can indeed see the shift in sentiment, but it’s not that no one is investing—at least for now, there’s quite a lot of investment interest in the market.

第四部分

登月第一步是“长上下文”——第二步是什么?

“接下来两个大里程碑”

张晓军:回到决定创业的时刻——你是一回国就启动了首轮融资吗?

杨植麟:从美国2月(去年)开始,部分是远程。最终国内投资者占多数。

张晓军:第一轮融了1亿美元吗?

杨植麟:第一轮没那么高;后面几轮超过了这个数字。2023年我们完成了两轮,总计近20亿人民币。

这是第三轮。我们尚未正式宣布融资,所以目前无法评论。

张晓军:有人说从2023年下半年起,没人愿意投资基础模型公司了。他们错了吗?

杨植麟:还是有的。你确实能看到情绪转变,但并非没人投资——至少目前,市场上有相当多的投资兴趣。

§ 19

Zhang Xiaojun: Besides capital and people, what other key decisions did you make in 2023?

Yang Zhilin: Deciding what to do. That’s the advantage of companies like ours—having a technical vision for decisions at the highest level.

We do long context. That requires judgment about the future: you need to know what is fundamental and where things are heading next. Again, it’s first principles—the process of “de-carving.” If you focus on carving, you can only look at what OpenAI has already done and figure out how to reproduce it.

You’ll find that doing lossless long-text compression in Kimi gives the product a unique experience. When you read English-language papers, it helps you understand them remarkably well. Using Claude or GPT-4 today, you won’t necessarily do as well; this required laying the groundwork in advance. We worked on it for over half a year. That’s very different from spotting a long-context trend today, hastily assembling two teams, and developing it at maximum speed.

Of course, the marathon has only just begun; more differentiation will come, and that requires you to anticipate in advance what counts as “a non-consensus that holds true.”

Zhang Xiaojun: In what month was this decision made?

Yang Zhilin: February or March—it was decided as soon as the company was founded.

Zhang Xiaojun: Why is long context the first step of the moonshot?

Yang Zhilin: Because it’s fundamental. It is the new computer’s memory.

The old computer’s memory grew by several orders of magnitude over the past few decades, and the same thing will happen with the new computer. It can solve many of today’s problems. For example, current multimodal architectures still need a tokenizer, but with a losslessly compressed long context, you don’t need one—you can put the raw input in directly. Taken further, it’s the foundation for making the new computing paradigm more general.

The old computer could represent everything with 0s and 1s; everything could be digitized. But today’s new computer can’t yet—there isn’t enough context, so it isn’t that general. To become a general world model, you need long context.

Second, it enables personalization. AI’s core value is personalized interaction; the value ultimately lands on personalization, and AGI will be more personalized than the previous generation of recommendation engines.

But personalization isn’t achieved through fine-tuning—it’s achieved by supporting very long context. Your entire history with the machine is context, and that context defines the personalization process. It cannot be replicated, and it makes for more direct dialogue—dialogue that generates information.

张晓军:除了资本和人,2023年你还做了哪些关键决策?

杨植麟:决定做什么。这是我们这类公司的优势——在最高决策层面拥有技术愿景。

我们做长上下文。这需要对未来的判断:你需要知道什么是根本的,以及下一步走向。又是第一性原理——即“去雕琢”的过程。如果你专注于雕琢,你只能看OpenAI已经做了什么,然后想办法复现。

你会发现,在Kimi中做无损长文本压缩给产品带来了独特体验。当你阅读英文论文时,它能很好地帮助你理解。用Claude或GPT-4,你未必能做得这么好;这需要提前打基础。我们在这方面做了超过半年。这与今天发现长上下文趋势、匆忙组建两个团队、以最快速度开发完全不同。

当然,马拉松才刚刚开始;更多的差异化会到来,这要求你提前预判什么算是“真实的非共识”。

张晓军:这个决策是几月做出的?

杨植麟:2月或3月——公司一成立就决定了。

张晓军:为什么长上下文是登月的第一步?

杨植麟:因为它是根本性的。它是新计算机的内存。

旧计算机的内存过去几十年增长了几个数量级,新计算机也会发生同样的事情。它可以解决当今的许多问题。例如,当前的多模态架构仍然需要分词器,但有了无损压缩的长上下文,你就不需要了——你可以直接输入原始数据。进一步说,它是让新计算范式更通用的基础。

旧计算机可以用0和1表示一切;一切都可以数字化。但今天的新计算机还不行——上下文不够,所以不够通用。要成为通用世界模型,你需要长上下文。

其次,它实现了个性化。AI的核心价值是个性化交互;价值最终落在个性化上,AGI将比上一代推荐引擎更个性化。

但个性化不是通过微调实现的——而是通过支持非常长的上下文实现。你与机器的整个历史就是上下文,这个上下文定义了个性化过程。它无法复制,并带来更直接的对话——产生信息的对话。

§ 20

Zhang Xiaojun: How much room is there to scale this up?

Yang Zhilin: Enormous. On one hand, expanding the window itself still has a long way to go—several orders of magnitude.

On the other hand, you can’t only expand the window, and you can’t just look at the number; whether the window is a few million tokens or several billion today is meaningless in itself. You have to look at the reasoning ability it enables within that window, the faithfulness—fidelity to the original information—and the instruction-following ability. You shouldn’t chase a single metric; you have to combine metrics with capabilities.

If these two dimensions keep improving, you can do a great deal. It could follow an instruction tens of thousands of words long, and the instruction itself could define many agents—highly personalized.

张晓军:扩展空间有多大?

杨植麟:巨大。一方面,扩展窗口本身还有很长的路要走——几个数量级。

另一方面,你不能只扩展窗口,也不能只看数字;窗口是几百万token还是几十亿,本身没有意义。你必须看它在窗口内实现的推理能力、忠实度——对原始信息的保真度——以及指令遵循能力。你不应该追求单一指标;你必须将指标与能力结合。

如果这两个维度持续提升,你可以做很多事。它可以遵循数万字的指令,指令本身可以定义许多代理——高度个性化。

§ 21

Zhang Xiaojun: Are the technologies behind long context and catching up with GPT-4 reusable for each other? Are they the same thing?

Yang Zhilin: I don’t think so. It’s more about adding a new dimension—a dimension GPT-4 doesn’t have.

Zhang Xiaojun: Many people say the leading Chinese foundation-model companies are all doing roughly the same thing—chasing GPT-3.5 in 2023, chasing GPT-4 in 2024. Do you agree?

Yang Zhilin: Improving general capabilities certainly has key milestones, so that statement is right to a degree—as a latecomer, you inevitably go through a catching-up process. But it’s also one-sided. Beyond general capabilities, there is a lot of space to develop distinctive capabilities and reach state-of-the-art in certain directions. Long context is one. DALL-E 3’s image generation is thoroughly outclassed by Midjourney V6. So you have to work on both fronts.

张晓军:长上下文和追赶GPT-4的技术能互相复用吗?是同一回事吗?

杨植麟:我认为不是。它更像是增加一个新维度——GPT-4没有的维度。

张晓军:许多人说中国领先的基础模型公司都在做差不多的事——2023年追赶GPT-3.5,2024年追赶GPT-4。你同意吗?

杨植麟:提升通用能力当然有关键里程碑,所以这个说法在一定程度上是对的——作为后来者,你不可避免地要经历追赶过程。但它也是片面的。除了通用能力,还有很大空间发展特色能力,在某些方向达到最先进水平。长上下文是一个。DALL-E 3的图像生成完全被Midjourney V6超越。所以你必须两条线作战。

§ 22

Zhang Xiaojun: If the moonshot’s first step is long context, what’s the second?

Yang Zhilin: There will be two big milestones ahead. First, a truly unified world model—one that unifies all the different modalities, a truly scalable and general architecture.

Second, enabling AI to keep evolving without human data input.

Zhang Xiaojun: How long will it take to reach these two milestones?

Yang Zhilin: Two to three years—possibly faster.

Zhang Xiaojun: So three years from now, we’ll already be looking at a world completely different from today’s.

Yang Zhilin: At the current pace of development, yes. The technology is now in a budding, fast-growing stage.

Zhang Xiaojun: Can you imagine what will exist three years from now?

Yang Zhilin: There will be a certain degree of AGI. Many of the things we do today, AI will also be able to do—even better than us. But the key is how we use it.

张晓军:如果登月第一步是长上下文,第二步是什么?

杨植麟:前方会有两个大里程碑。首先,一个真正统一的世界模型——统一所有不同模态,一个真正可扩展且通用的架构。

其次,让AI能够无需人类数据输入而持续进化。

张晓军:需要多久才能达到这两个里程碑?

杨植麟:两到三年——可能更快。

张晓军:所以三年后,我们将看到一个与今天完全不同的世界。

杨植麟:按照目前的发展速度,是的。技术正处于萌芽和快速增长阶段。

张晓军:你能想象三年后会有什么吗?

杨植麟:会有一定程度的AGI。我们今天做的许多事情,AI也能做——甚至比我们更好。但关键是我们如何使用它。

§ 23

Zhang Xiaojun: Will you go all in on catching up with GPT-4?

Yang Zhilin: GPT-4 is a necessary stop on the road to AGI. The key is not to be satisfied with merely matching GPT-4. First, you have to ask what the real non-consensus is now: beyond GPT-4, what’s next? What should GPT-5 and GPT-6 look like? Second, you have to see which distinctive capabilities you have within that—and that matters more.

张晓军:你会全力追赶GPT-4吗?

杨植麟:GPT-4是通往AGI的必经站点。关键是不能满足于仅仅追上GPT-4。首先,你要问现在真正的非共识是什么:超越GPT-4,下一步是什么?GPT-5和GPT-6应该长什么样?其次,你要看到其中你有什么特色能力——这更重要。

§ 24

Zhang Xiaojun: Other foundation-model companies publish their model capabilities and rankings. You don’t seem to have done that?

Yang Zhilin: Chasing leaderboard rankings means very little. The best leaderboard is the users—you should let users vote. Many leaderboards have problems.

Zhang Xiaojun: Is being the fastest to reach GPT-4 among China’s foundation-model companies your goal? Does fast versus slow make a difference?

Yang Zhilin: Definitely. Over a long enough time frame, everyone will eventually get there. But it depends on how long your lead or lag is. A gap of six months or more is meaningful—and it also depends on what you can do with that window.

Zhang Xiaojun: When do you expect to reach GPT-4?

Yang Zhilin: It should be quite soon, but I can’t disclose the specific timing publicly.

Zhang Xiaojun: Will you be the fastest?

Yang Zhilin: That has to be assessed dynamically—but we have a real chance.

Zhang Xiaojun: After launching Kimi, what is your North Star metric?

Yang Zhilin: Today it’s about making the product better and adding more dimensions. For example, we shouldn’t just be fighting tooth and nail over a search use case—search will later be only a small fraction of this product’s value; the product should have a much bigger increment. Being 10% or 20% better than a traditional search engine isn’t worth much—only something truly disruptive deserves the three letters “AGI.”

The unique value is your incremental intelligence. You have to hold onto this point: intelligence is always the core incremental value. If only 10%–20% of your product’s core value comes from AI, it doesn’t hold up.

张晓军:其他基础模型公司公布模型能力和排名。你们似乎没有这么做?

杨植麟:追逐排行榜意义不大。最好的排行榜是用户——你应该让用户投票。很多排行榜有问题。

张晓军:成为中国基础模型公司中最快达到GPT-4的是你的目标吗?快慢有区别吗?

杨植麟:当然。在足够长的时间框架内,最终大家都会达到。但这取决于你领先或落后多长时间。六个月的差距是有意义的——还取决于你能利用那个窗口做什么。

张晓军:你预计何时达到GPT-4?

杨植麟:应该很快,但我不能公开透露具体时间。

张晓军:你会是最快的吗?

杨植麟:这需要动态评估——但我们确实有机会。

张晓军:推出Kimi后,你的北极星指标是什么?

杨植麟:目前是让产品更好,增加更多维度。例如,我们不应该只在一个搜索用例上死磕——搜索以后只是这个产品价值的一小部分;产品应该有更大的增量。比传统搜索引擎好10%或20%没什么价值——只有真正颠覆性的东西才配得上“AGI”这三个字母。

独特价值是你的智能增量。你必须坚持这一点:智能永远是核心增量价值。如果你的产品核心价值只有10%-20%来自AI,那站不住脚。

§ 25

Part 5

I’m Not at All Anxious About Commercialization

“User Scaling and Model Scaling Need to Happen at the Same Time”

Zhang Xiaojun: Mid-2023 was a huge watershed—the market turned from frenzy to chill very quickly. How did you perceive it?

Yang Zhilin: I don’t fully agree with that characterization—we did complete a funding round in the second half of the year. And new things kept coming out. Today’s model capabilities were unimaginable at the end of last year. The user numbers and revenue of more and more AI companies kept rising. It has continuously proven its value.

Zhang Xiaojun: For you, what felt different between the first and second halves of the year?

Yang Zhilin: Not much changed. Variables certainly exist, but you return to first principles—how to give users a good product. Ultimately, we must satisfy user needs, not win a race. We are not a company built for competition.

Zhang Xiaojun: The industry believes a notable difference between the first and second halves of 2023 was a shift of focus: the first half was more about AGI; the second half turned to how to land applications and commercialize. Did you make that shift?

Yang Zhilin: Of course I’m going to do AGI—it’s the only meaningful thing for the next decade. But that doesn’t mean we don’t build applications. Or rather, it shouldn’t be defined as an “application.”

“Application” makes it sound like you have a technology and you’re looking for somewhere to use it, with a commercial loop closed. But “application” isn’t the accurate word. It and AGI complement each other. It is both the means to achieve AGI and the purpose of achieving it. “Application” sounds more like a purpose: I want to make it useful. You have to combine Eastern and Western philosophies—you have to make money, and you have to have ideals.

Today, users help us discover scenarios we never considered. Someone uses it to screen résumés—something we never thought of when designing the product, but it naturally works. User input, in turn, makes the model better. Why is Midjourney so good? It scaled on the user side—user scaling and model scaling must happen at the same time. Conversely, if you only focus on applications and ignore the iteration of model capabilities—ignore AGI—your contribution will be limited.

第五部分

我对商业化一点也不焦虑

“用户扩展和模型扩展必须同时发生”

张晓军:2023年中是一个大分水岭——市场从狂热迅速转为冷淡。你怎么看?

杨植麟:我不完全同意这种说法——我们在下半年确实完成了一轮融资。而且新东西不断涌现。今天模型的能力在去年底是无法想象的。越来越多AI公司的用户和收入持续增长。它不断证明自己的价值。

张晓军:对你来说,上半年和下半年感觉有什么不同?

杨植麟:变化不大。变量当然存在,但你回归第一性原理——如何给用户一个好的产品。最终我们必须满足用户需求,而不是赢得竞赛。我们不是为竞争而建立的公司。

张晓军:业界认为2023年上下半年的一个显著差异是焦点转移:上半年更关注AGI,下半年转向如何落地应用和商业化。你做了这个转变吗?

杨植麟:我当然要做AGI——这是未来十年唯一有意义的事情。但这并不意味着我们不构建应用。或者说,不应该将其定义为“应用”。

“应用”听起来好像你有一项技术,然后寻找地方使用它,形成一个商业闭环。但“应用”这个词不准确。它和AGI相互补充。它既是实现AGI的手段,也是实现AGI的目的。“应用”听起来更像一个目的:我想让它有用。你必须结合东方和西方哲学——既要赚钱,又要有理想。

如今,用户帮助我们发现了我们从未考虑过的场景。有人用它筛选简历——我们在设计产品时从未想过,但它自然就奏效了。用户的输入反过来让模型变得更好。为什么Midjourney这么好?它在用户侧扩展了——用户扩展和模型扩展必须同时发生。反之,如果你只关注应用,忽略模型能力的迭代——忽略AGI——你的贡献将是有限的。

§ 26

Zhang Xiaojun: Zhu Xiaohu, managing partner at GSR Ventures, only invests in foundation-model applications. One of his views: the hardest core problem is PMF for AIGC—if ten people can’t find PMF, a hundred people won’t either; it has nothing to do with headcount or cost, so don’t burn money on it. He says, “Train on LLaMA for two or three months and you can at least reach the level of the top 30 humans—it can replace people immediately.” What do you think of his view?

Yang Zhilin: AI is not about what PMF I can find in the next year or two; it’s about how to change the world over the next ten to twenty years—these are two different ways of thinking.

We are staunch long-termists. When AGI or something stronger is achieved, everything today will be rewritten. PMF is certainly important, but if you rush to find PMF, you’ll very likely be hit by another “dimensionality-reduction strike”—being crushed by a higher-dimensional technology. That has happened too many times. In the past, many people built customer-service and dialogue systems, doing slot filling—some were companies of decent scale. But they were all wiped out by a higher-dimensional blow. It was painful.

That’s not to say the approach never works. Suppose you find a scenario today where current technology suffices, where the 0-to-1 incremental value is enormous and the 1-to-n space isn’t that big—that scenario is fine. Midjourney is like that, or copywriting generation—relatively simple tasks with very visible 0-to-1 effects. Those are opportunities for the applications-only camp. But the biggest opportunity isn’t there. If your premise is commercialization, you can’t think about it apart from AGI. If I only build applications now—fine, but in a year you could be crushed.

Zhang Xiaojun: You could quietly upgrade the underlying model, couldn’t you?

Yang Zhilin: But that approach can never become bigger than the model itself. Technology is the only new variable of this era; the other variables haven’t changed. Returning to first principles, AGI is the core of everything. From that, we deduced: a super app definitely requires the strongest technical capabilities.

张晓军:金沙江创投管理合伙人朱啸虎只投基础模型应用。他的一个观点:AIGC最难的核心问题是PMF——如果十个人找不到PMF,一百个人也找不到;与人数或成本无关,所以别烧钱。他说:“用LLaMA训练两三个月,至少能达到人类前30名的水平——可以立即替代人。”你怎么看?

杨植麟:AI不是关于我未来一两年能找到什么PMF,而是关于未来十到二十年如何改变世界——这是两种不同的思维方式。

我们坚定的长期主义者。当AGI或更强的东西实现时,今天的一切都将被重写。PMF当然重要,但如果你急于找到PMF,很可能会被另一个“降维打击”——被更高维度的技术压垮。这已经发生过太多次了。过去,很多人做客服和对话系统,做槽位填充——有些公司规模还不错。但都被更高维度的打击消灭了。很痛苦。

这并非说这种方法永远无效。假设你今天找到一个场景,现有技术足够,0到1的增量价值巨大,1到n的空间不是很大——这个场景没问题。Midjourney就是这样,或者文案生成——相对简单的任务,0到1效果非常明显。这些是纯应用派的机会。但最大的机会不在这里。如果你的前提是商业化,你不能脱离AGI来思考。如果我现在只构建应用——可以,但一年后你可能被压垮。

张晓军:你可以偷偷升级底层模型,不是吗?

杨植麟:但那种方法永远无法比模型本身更大。技术是时代唯一的新变量;其他变量没有改变。回到第一性原理,AGI是一切的核心。由此我们推导出:超级应用绝对需要最强的技术能力。

§ 27

Zhang Xiaojun: Can you use open-source models? (The latest news is that Google announced the open-source model Gemma.)

Yang Zhilin: Open source lags behind closed source—that’s also a fact.

Zhang Xiaojun: Might the lag be only temporary?

Yang Zhilin: It doesn’t look that way so far.

Zhang Xiaojun: Why can’t open source catch up with closed source?

Yang Zhilin: Because open-source development works differently now. In the past, everyone could contribute to open source; today, open source itself is still centralized. Many open-source contributions probably haven’t been validated by compute. Closed source enjoys concentrations of talent and capital, so in the end closed source will definitely be better—it’s a consolidation.

If I had a leading model today, open-sourcing it would very likely be irrational. It’s the laggards who might do that instead, or open-source a small model—to stir things up; after all, if you don’t open-source it, it has no value anyway.

张晓军:你们可以用开源模型吗?(最新消息是谷歌宣布开源模型Gemma。)

杨植麟:开源落后于闭源——这也是事实。

张晓军:这种落后可能只是暂时的吗?

杨植麟:目前看来并非如此。

张晓军:为什么开源追不上闭源?

杨植麟:因为开源的发展方式已经不同。过去,每个人都可以为开源做贡献;今天,开源本身仍然是中心化的。许多开源贡献可能尚未经过算力验证。闭源拥有人才和资本的集中,因此最终闭源肯定更好——这是一种整合。

如果我现在拥有领先模型,将其开源很可能是非理性的。反而是落后者可能会这样做,或开源一个小模型——以搅动局面;毕竟,不开源它也没有价值。

§ 28

Zhang Xiaojun: How do you push back against the anxiety in China? People say that a foundation-model company that doesn’t quickly produce commercial scenarios and products that meet investor expectations will struggle to raise its next round.

Yang Zhilin: You need a balance between the long term and the short term. Having no users and no revenue at all definitely won’t work.

As we’ve seen, going from GPT-3.5 to GPT-4 unlocked many applications; from GPT-4 to GPT-4.5 and then GPT-5, it will very likely keep unlocking more—even exponentially more. The so-called “Moore’s law of scenarios” means the number of usable scenarios rises exponentially over time. We need to improve model capabilities while finding more scenarios—that kind of balance.

It’s a spiral. It depends on how much of your investment goes to the short term and how much to the long term. You pursue the long term on the condition that you can survive. The long term absolutely cannot be abandoned, or you’ll miss the entire era. Drawing conclusions today is truly too early.

张晓军:你如何应对中国的焦虑?有人说,不能快速产生符合投资者期望的商业场景和产品的基础模型公司将难以进行下一轮融资。

杨植麟:你需要在长期和短期之间取得平衡。完全没有用户和营收肯定不行。

正如我们所见,从GPT-3.5到GPT-4解锁了许多应用;从GPT-4到GPT-4.5再到GPT-5,很可能会解锁更多——甚至指数级增加。所谓的“场景摩尔定律”意味着可用场景的数量随时间指数增长。我们需要在提升模型能力的同时找到更多场景——这种平衡。

这是一个螺旋。它取决于你的投资有多少用于短期,多少用于长期。你在生存的基础上追求长期。长期绝对不能被放弃,否则你会错过整个时代。今天下结论确实太早了。

§ 29

Zhang Xiaojun: Do you agree with the “two-wheel drive” idea put forward by Wang Huiwen, co-founder of Meituan and founder of Light Year?

Yang Zhilin: That’s a good question. To a degree, the logic holds. But how you actually execute makes a huge difference. Can you truly make some “non-consensus bets with favorable odds”?

Zhang Xiaojun: As I understand it, their two-wheel drive also requires quickly finding new application scenarios; otherwise, there’s no way for the technology to land.

Yang Zhilin: It still comes down to the difference between model scaling and user scaling.

Zhang Xiaojun: In China, besides you, who else takes the model-scaling approach?

Yang Zhilin: That’s not for me to judge.

Zhang Xiaojun: Most people probably take the user-scaling approach. Or can we put it this way: is this the difference between the academic camp and the commercialization camp?

Yang Zhilin: We are not the academic camp. The academic camp definitely doesn’t work.

Zhang Xiaojun: Many foundation-model companies commercialize through B2B—after all, B2B offers more certainty. Do you?

Yang Zhilin: We don’t. From day one, we decided to go B2C.

It depends on what you want. If you know something isn’t what you want, you won’t get FOMO—because even if you got it, it wouldn’t mean much.

张晓军:你同意美团联合创始人、光年之外创始人王慧文提出的“双轮驱动”想法吗?

杨植麟:好问题。在某种程度上,逻辑成立。但实际执行方式差异很大。你能否真正做出一些“赔率有利的非共识赌注”?

张晓军:据我理解,他们的双轮驱动也需要快速找到新应用场景;否则技术无法落地。

杨植麟:归根结底还是模型扩展和用户扩展的区别。

张晓军:在中国,除了你,还有谁采取模型扩展路线?

杨植麟:这个我不评判。

张晓军:大多数人可能采取用户扩展路线。或者可以这么说:这是学术派和商业化派的区别吗?

杨植麟:我们不是学术派。学术派肯定不行。

张晓军:许多基础模型公司通过B2B商业化——毕竟B2B更确定。你们呢?

杨植麟:我们不。从第一天起,我们就决定做B2C。

这取决于你想要什么。如果你知道某件事不是你想要的东西,你就不会FOMO——因为即使得到了,也没什么意义。

§ 30

Zhang Xiaojun: Have you been anxious over the past year?

Yang Zhilin: More excitement and exhilaration. Because I’ve thought about this for a very long time. We were probably among the earliest who wanted to explore the dark side of the moon. Today you find that you’re really building a rocket, and every day you’re discussing what fuel to add to make it go faster—and how to keep it from blowing up.

Zhang Xiaojun: To sum up the “non-consensus bets with favorable odds” you’ve made—besides B2C and long context, are there others?

Yang Zhilin: More are in the works; I hope to share them with everyone soon.

Zhang Xiaojun: China’s previous generation of entrepreneurs tasted success with applications and scenarios, so they focus more on product, users, and the data flywheel. Can the new generation of AI entrepreneurs you represent stand for a new future?

Yang Zhilin: We care deeply about users, too. Users are our ultimate goal, but it’s also a process of co-creation. The biggest difference is that this time it will be more technology-driven. It’s still the horse-carriage-versus-car question: we’re now in the leap from horse carriages to cars, and we should focus as much as possible on how to give users a car.

Zhang Xiaojun: Do you feel lonely?

Yang Zhilin: Ha ha ha… That’s an interesting question. I think I’m fine, because I still have dozens—nearly a hundred—people fighting alongside me.

张晓军:过去一年你焦虑过吗?

杨植麟:更多的是兴奋和激动。因为我想这件事想了很久。我们可能是最早想探索月之暗面的人之一。今天你发现你真的在造火箭,每天都在讨论加什么燃料让它飞得更快——以及如何防止它爆炸。

张晓军:总结一下你做出的“赔率有利的非共识赌注”——除了B2C和长上下文,还有吗?

杨植麟:更多在进行中;希望很快能和大家分享。

张晓军:中国上一代创业者通过应用和场景尝到甜头,所以他们更关注产品、用户和数据飞轮。你代表的新一代AI创业者能否代表新的未来?

杨植麟:我们也非常关心用户。用户是我们的最终目标,但也是一个共创的过程。最大的不同是,这次会更加技术驱动。还是马车与汽车的问题:我们现在正处于从马车到汽车的飞跃,我们应该尽可能专注于如何给用户一辆汽车。

张晓军:你感到孤独吗?

杨植麟:哈哈哈……这是个有趣的问题。我觉得还好,因为我还有几十个——将近一百人——和我并肩作战。

§ 31

Part 6

Before We’ve Even Caught up with GPT-4, Sora Arrives

“Right Now It’s Like the GPT-3.5 Moment for Video Generation”

Zhang Xiaojun: Sora’s sudden appearance this year—how much of it was within your expectations and how much was not?

Yang Zhilin: That generative AI could achieve this effect was within expectations; what was unexpected was the timing—it came earlier than we had estimated. It also reflects how fast AI is developing now: a lot of the dividends of scaling haven’t been fully harvested yet.

Zhang Xiaojun: Last year, the industry judged that foundation models in 2024 would inevitably compete hard on multimodal narratives, and that video generation quality would improve as rapidly as text-to-image did in 2023. Did Sora’s technical capabilities exceed, meet, or fall short of your expectations?

Yang Zhilin: It solved many previously difficult problems. For example, maintaining consistency of generation over a relatively long time window—that’s the key point, and a huge improvement.

Zhang Xiaojun: What does it mean for the global industry landscape? What new narratives will foundation models see in 2024?

Yang Zhilin: First, near-term application value: it can further improve efficiency in production processes, and of course we hope for more extensions built on current capabilities. Second, combining with other modalities. It is itself a model of the world; with that knowledge, it’s an excellent complement to existing text. On that basis, there is quite a lot of room and opportunity, whether in agents or in connecting with the physical world.

第六部分

还没追上GPT-4,Sora来了

“现在就像视频生成的GPT-3.5时刻”

张晓军:今年Sora的突然出现——有多少在你意料之中,多少出乎意料?

杨植麟:生成式AI能达到这个效果是在意料之中的;出乎意料的是时间——它比我们预估的来得早。这也反映了AI现在发展有多快:规模化的很多红利还没被充分收割。

张晓军:去年业界判断,2024年基础模型必然会在多模态叙事上激烈竞争,视频生成质量会像2023年文生图一样快速提升。Sora的技术能力是超出、符合还是低于你的预期?

杨植麟:它解决了很多以前困难的问题。比如,在相对长的时间窗口内保持生成一致性——这是关键点,一个巨大的进步。

张晓军:对全球产业格局意味着什么?2024年基础模型会有哪些新叙事?

杨植麟:首先,近期应用价值:它可以进一步提高生产流程的效率,当然我们希望有更多基于当前能力的扩展。其次,与其他模态结合。它本身就是一个世界模型;有了这个知识,它与现有文本是极好的互补。在此基础上,有相当多的空间和机会,无论是在代理方面还是在与物理世界的连接方面。

§ 32

Zhang Xiaojun: Overall, how do you assess Sora?

Yang Zhilin: We had been planning a similar direction ourselves and had worked on it for a while. Directionally, it wasn’t much of a surprise—it’s more about the technical details.

Zhang Xiaojun: What technical details are worth learning from?

Yang Zhilin: OpenAI hasn’t fully explained many of them either. They described the broad strokes; there are key details you have to infer from its outputs or from available information, combined with our own earlier experiments. At least for us, it adds more data points to the development process—more data input.

Zhang Xiaojun: Compared with text generation, what were the main bottlenecks in video generation? What solutions can you see OpenAI found this time?

Yang Zhilin: The main bottleneck—the core is still data: how do you fit the data at scale? That hadn’t been validated before, especially when the motion is complex and the generated result is photorealistic. Under those conditions, being able to scale—that’s what it solved this time.

What remains unsolved includes, for example, the need for a unified architecture. DiT is still not a very general architecture. Modeling the marginal probability of purely visual signals can be done very well, but how do you generalize that into a universal new computer? You still need a more unified architecture—there’s still room there.

Zhang Xiaojun: Have you read OpenAI’s Sora report, “Video generation models as world simulators”? What key points in it deserve highlighting?

Yang Zhilin: I have. Given the current competitive situation, they definitely wouldn’t write down the most important points. But it’s still worth learning from. This was essentially paid content—things you might otherwise have to spend money on many experiments to learn. Now you can know some of it without paying for the experiments, and form a rough understanding.

Zhang Xiaojun: What key signals did you extract from it?

Yang Zhilin: That this thing is scalable to a degree. In addition, it gives a fairly concrete account of how the architecture is built. But it’s also possible that different architectures don’t make such an essential difference on this problem.

Zhang Xiaojun: Do you agree with that line of theirs—“Scaling video generation models is a promising path towards building general-purpose simulators of the physical world”?

Yang Zhilin: I strongly agree. These two things optimize the same objective function—there’s not much doubt about that.

Zhang Xiaojun: What do you think of Yann LeCun once again speaking out against generative AI? His view: “Modeling the world by generating pixels is wasteful and doomed to fail. Generation happens to work for text because text is discrete, with a finite number of symbols. In that case, dealing with uncertainty in prediction is easy; dealing with predictive uncertainty in high-dimensional continuous sensory inputs is intractable.”

Yang Zhilin: I now think that when you model the marginal probability of video, the essence is lossless compression—no essential difference from next-token prediction in language models. As long as you compress well enough, you can explain whatever in this world is explainable.

But there’s also important work yet to be done: how does it combine with the capabilities that have already been compressed? You can think of it as two different kinds of compression. One compresses the raw world—that’s what video models do. The other compresses the behaviors humans produce, because human behavior has passed through the human brain—the only thing in the world that produces intelligence. You can think of video models as doing the first kind and text models the second, though video models also contain some of the second kind: some human-made videos contain the intelligence of their creators. Ultimately, it will probably be a mix—you need to learn from different angles through both approaches, and both contribute to the growth of intelligence.

So generation may not be the goal; it is merely the compression function. If you compress well enough, the generation will be very good in the end. Conversely, if a model itself cannot generate, is it still possible to compress extremely well? That’s doubtful. It’s possible that generating very well is a necessary condition for compressing very well.

Zhang Xiaojun: Sora and last year’s ChatGPT are two different milestones. Which is bigger?

Yang Zhilin: Both are very important. Right now it’s a bit like the GPT-3.5 moment for video generation—a step-function improvement. The model is still relatively small, and it’s foreseeable that there will be bigger models, which is a guaranteed improvement in capability.

张晓军:总体而言,你如何评估Sora?

杨植麟:我们自己也在规划类似方向,并已经做了一段时间。方向上并不意外——更多是技术细节问题。

张晓军:哪些技术细节值得学习?

杨植麟:OpenAI很多也没有完全解释。他们描述了框架;关键细节需要从输出或可用信息中推断,结合我们自己的早期实验。至少对我们来说,它给开发过程增加了更多数据点——更多数据输入。

张晓军:与文本生成相比,视频生成的主要瓶颈是什么?你觉得OpenAI这次找到了什么解决方案?

杨植麟:主要瓶颈——核心还是数据:如何大规模适配数据?以前没有验证过,尤其是当运动复杂且生成结果逼真时。在这种情况下,能够规模化——这就是这次解决的问题。

尚未解决的问题包括,例如,需要统一架构。DiT仍然不是一个非常通用的架构。对纯视觉信号的边际概率建模可以做得很好,但如何将其推广到通用新计算机?仍然需要一个更统一的架构——那里还有空间。

张晓军:你读过OpenAI的Sora报告《视频生成模型作为世界模拟器》吗?有哪些关键点值得强调?

杨植麟:读过。鉴于当前的竞争形势,他们肯定不会写下最重要的点。但仍然值得学习。这基本上是付费内容——否则你可能需要花很多实验才能学到的东西。现在你无需为实验付费就能了解一部分,形成粗略理解。

张晓军:你从中提取了什么关键信号?

杨植麟:这个东西在某种程度上是可扩展的。此外,它对架构的构建给出了相当具体的描述。但也有可能不同的架构在这个问题上不会产生本质差异。

张晓军:你同意他们那句话吗——“扩展视频生成模型是构建物理世界通用模拟器的有前途路径”?

杨植麟:我强烈同意。这两者优化的是同一个目标函数——这点没什么疑问。

张晓军:你怎么看Yann LeCun再次发声反对生成式AI?他的观点:“通过生成像素来建模世界是浪费且注定失败的。生成之所以对文本有效,是因为文本是离散的,符号数量有限。在这种情况下,处理预测中的不确定性很容易;而在高维连续感官输入中处理预测不确定性是难以处理的。”

杨植麟:我现在认为,当你对视频的边际概率建模时,本质是无损压缩——与语言模型中的下一个token预测没有本质区别。只要你压缩得足够好,这个世界上任何可解释的东西都能解释。

但还有重要工作要做:如何与已经压缩的能力结合?你可以将其视为两种不同的压缩。一种压缩原始世界——这是视频模型做的。另一种压缩人类产生的行为,因为人类行为经过了人脑——世界上唯一产生智能的东西。你可以认为视频模型做第一种,文本模型做第二种,尽管视频模型也包含一些第二种:一些人造视频包含了创作者的智能。最终,很可能是混合——你需要通过两种方法从不同角度学习,两者都对智能增长有贡献。

所以生成可能不是目标;它只是压缩函数。如果你压缩得足够好,最终生成会非常好。反之,如果一个模型本身不能生成,它还能压缩得非常好吗?值得怀疑。很可能生成得非常好是压缩得非常好的必要条件。

张晓军:Sora和去年的ChatGPT是两个不同的里程碑。哪个更大?

杨植麟:两者都非常重要。现在有点像视频生成的GPT-3.5时刻——一个阶跃改进。模型仍然相对较小,可以预见会有更大的模型,这是能力的保证提升。

§ 33

Zhang Xiaojun: Some people also say that, for multimodality, Google Gemini’s breakthrough matters more.

Yang Zhilin: Gemini follows the GPT-4V line and incorporates that understanding as well. Both are important; the final step is putting all these things into the same model, and that hasn’t been solved yet.

Zhang Xiaojun: Why is putting them in the same model so hard?

Yang Zhilin: Nobody knows how yet. There is still no validated architecture.

Zhang Xiaojun: What will Sora + GPT produce?

Yang Zhilin: Sora can be applied to video production right away, but if combined with language models, it could connect the digital world and the physical world. You could also complete tasks more end-to-end, because your modeling of the world is now better than before—it can even be used to improve your understanding of multimodal inputs. So in the end you can switch quite fluidly between modalities.

To sum up: you understand the world better; you can do more end-to-end tasks in the digital world; and you can even build a bridge to the physical world to complete tasks there. This is the starting point. Autonomous driving, for example, or some household chores—these are, in theory, all instances of connecting with the physical world. So the breakthrough in the digital world is certain, but it also holds the potential of a path to the physical.

Zhang Xiaojun: What does Sora mean for Chinese foundation-model companies? What’s the right response?

Yang Zhilin: It doesn’t change much. This was always a direction of certainty.

Zhang Xiaojun: Chinese foundation models haven’t caught up with GPT-4 yet, and now Sora has arrived. What do you think? The two worlds seem to be drifting further and further apart—do you feel anxious?

Yang Zhilin: It’s simply an objective fact. But the actual gap may still be narrowing—that’s the law of technological development.

Zhang Xiaojun: What do you mean? That the technology curve is steep at first and then gradually flattens?

Yang Zhilin: Yes. I’m not really surprised—OpenAI has been working on next-generation models all along. Objectively, the gap will persist for a while, and gaps between different Chinese companies will persist for a while too—this is a period of technological explosion.

But in another two or three years, it’s possible that China’s top companies can do more of the foundational work well here—including technical infrastructure, talent reserves, and the sedimentation of organizational culture. With that honing, they’ll be more likely to lead in certain respects—but it will take some patience.

Zhang Xiaojun: Could China and the U.S. end up with completely different AI technology ecosystems?

Yang Zhilin: The ecosystems could differ, if you look at it from a product and commercialization angle. But from a technology angle, general capabilities won’t follow completely different technical routes—the basic general capabilities will definitely be similar. Because the space of AGI is vast, differentiation on top of general capabilities is more likely to happen.

张晓军:也有人说,在多模态方面,谷歌Gemini的突破更重要。

杨植麟:Gemini延续GPT-4V路线,也包含了那种理解。两者都很重要;最后一步是把所有这些放到同一个模型中,这一点尚未解决。

张晓军:为什么把它们放在同一个模型里这么难?

杨植麟:还没有人知道怎么做。目前还没有经过验证的架构。

张晓军:Sora + GPT会产生什么?

杨植麟:Sora可以立即应用于视频制作,但如果与语言模型结合,它可以连接数字世界和物理世界。你还可以更端到端地完成任务,因为你现在对世界的建模比以前更好——它甚至可以用来改善你对多模态输入的理解。所以最终你可以在模态之间非常流畅地切换。

总结:你更好地理解世界;你可以在数字世界中做更多端到端的任务;你甚至可以搭建一座通往物理世界的桥梁,在那里完成任务。这是起点。例如自动驾驶,或一些家务——这些理论上都是与物理世界连接的实例。所以数字世界的突破是确定的,但它也蕴藏着通往物理世界的路径潜力。

张晓军:Sora对中国基础模型公司意味着什么?正确的应对方式是什么?

杨植麟:变化不大。这本来就是一个确定的方向。

张晓军:中国基础模型还没追上GPT-4,现在Sora又来了。你怎么看?两个世界似乎越走越远——你焦虑吗?

杨植麟:这只是一个客观事实。但实际差距可能仍在缩小——这是技术发展的规律。

张晓军:你的意思是技术曲线起初陡峭,然后逐渐变平?

杨植麟:是的。我并不特别惊讶——OpenAI一直在做下一代模型。客观上说,差距还会持续一段时间,中国不同公司之间的差距也会持续一段时间——这是一个技术爆炸期。

但再过两三年,中国顶尖公司有可能在这里做好更多基础工作——包括技术基础设施、人才储备、组织文化的沉淀。通过这些磨练,它们更有可能在某些方面领先——但这需要一点耐心。

张晓军:中美最终会不会有完全不同的AI技术生态?

杨植麟:从产品和商业化角度看,生态可能不同。但从技术角度看,通用能力不会走完全不同的技术路线——基本通用能力肯定会类似。因为AGI的空间巨大,在通用能力之上更容易产生差异化。

§ 34

Zhang Xiaojun: There’s a long-running debate in Silicon Valley: “one model rules all” versus “many specialized (smaller) models”—one general model for all kinds of tasks, or many specialized smaller models for specific tasks. What’s your view?

Yang Zhilin: I take the first view.

Zhang Xiaojun: On this point, will China and the U.S. diverge greatly?

Yang Zhilin: I don’t think so, ultimately.

张晓军:硅谷有一个长期争论:“一个模型统治一切” vs “许多专用(更小)模型”——一个通用模型处理所有任务,或许多专用小模型处理特定任务。你怎么看?

杨植麟:我持第一种观点。

张晓军:在这一点上,中美会有很大分歧吗?

杨植麟:我认为最终不会。

§ 35

Part 7

I Accept That Failure Is a Possibility

“It Has Already Changed My Life”

Zhang Xiaojun: Foundation-model entrepreneurship is a rather peculiar creature in China: you’ve raised so much money, yet it seems a large chunk of it goes toward scientific experiments. How do you persuade investors to open their wallets under these circumstances?

Yang Zhilin: No differently than in the U.S. The money we’ve raised today isn’t even that much. So we still have a lot to learn from OpenAI.

Zhang Xiaojun: I’d like to know: how much more money does it take to reach GPT-4? How much to reach Sora?

Yang Zhilin: Neither GPT-4 nor Sora requires that much. The money now is more about reserving for the next generation—or the generation after that—of models, for frontier exploration.

Zhang Xiaojun: Chinese foundation-model startups have taken the giants’ money, but the giants are also training their own models. How do you view the relationship between foundation-model startups and the giants?

Yang Zhilin: There’s both competition and cooperation. The giants and the startups have different first priorities. Look at any big tech company today: its first priority differs from an AGI company’s first priority. Your first priority shapes your actions and results, and ultimately defines the different relationships within the ecosystem.

Zhang Xiaojun: Why do the giants spread smaller investments across multiple foundation-model companies instead of betting heavily on one?

Yang Zhilin: It’s a matter of stage. There will be more consolidation going forward—and fewer companies.

Zhang Xiaojun: Some say the endgame for foundation-model companies is being acquired by a giant. Do you agree?

Yang Zhilin: Not necessarily, I think. But they may well have very deep partnerships.

Zhang Xiaojun: For example, how might they cooperate?

Yang Zhilin: OpenAI and Microsoft are the classic model of cooperation. Much of it can be referenced, and some of it can be improved.

Zhang Xiaojun: Over the past year, where did the twists and turns of entrepreneurship show up for you?

Yang Zhilin: There were many external variables—capital, talent, GPUs, product, R&D, technology. There were highlight moments, and there were difficulties to overcome. Take GPUs.

There was a lot of back and forth: supply was tight for a while, then it improved. The most extreme stretch saw prices change daily—a machine might cost 260 one day, 340 the next, then fall back two days later. It was a constantly moving situation. You had to watch it closely. Prices kept changing, so strategy kept changing: which channels to use, whether to buy or rent—there were many different options.

Zhang Xiaojun: What drove these fluctuations?

Yang Zhilin: Geopolitical reasons; production itself comes in batches; and market sentiment plays a role. We observed many companies starting to return GPUs, realizing they didn’t necessarily need to train this model. As market sentiment and people’s decisions shifted, supply and demand shifted with them. The good news is that overall supply has improved enormously lately. My personal judgment is that for at least the next one to two years, GPUs won’t be a major bottleneck.

Zhang Xiaojun: You seem to think constantly about organization. How have you gone about team building?

Yang Zhilin: Our approach to hiring evolved. AGI talent is extremely limited worldwide, and people with relevant experience are rare. Our earliest hiring profile focused on finding geniuses with directly relevant skills. That proved very successful. People who had previously “performed surgery” on models and had firsthand experience training ultra-large-scale models could build things very quickly. The Kimi launch included—the capital efficiency and organizational efficiency were actually very high.

Zhang Xiaojun: How much did it cost?

Yang Zhilin: A rather small number—compared with many other expenses, it was doing big things with small money. For a long time we were at 30–40 people. Now we’re at 80. We pursue talent density.

The talent profile changed later. In the earliest days we hired geniuses, believing their ceiling was high—a company’s ceiling is determined by the ceiling of its people. Later we rounded out the team with more dimensions of talent—people on the product-operations side, leader types, people who can take things to the extreme. Now it’s a more complete, resilient team that can fight.

Zhang Xiaojun: After a year of foundation-model entrepreneurship in China, how do you assess the milestone results so far?

Yang Zhilin: We’ve built a rocket prototype and are now test-firing it. We’ve assembled a team, figured out some of the fuel formulas, and can more or less see an embryonic PMF. You could say we’ve taken the first step of the moonshot.

第七部分

我接受失败是可能的

“它已经改变了我的人生”

张晓军:基础模型创业在中国是个相当奇特的存在:你融了那么多钱,但看起来很大一部分用在了科学实验上。在这种情况下,你怎么说服投资者掏钱?

杨植麟:和在美国没什么不同。我们今天融的钱甚至不算多。所以我们还有很多地方要向OpenAI学习。

张晓军:我想知道:要达到GPT-4需要多少钱?要达到Sora呢?

杨植麟:GPT-4和Sora都不需要那么多。现在的钱更多是为下一代——或下下一代——模型预留,用于前沿探索。

张晓军:中国基础模型初创公司拿了巨头的钱,但巨头自己也在训练模型。你怎么看基础模型初创公司与巨头的关系?

杨植麟:既有竞争也有合作。巨头和初创公司的首要优先事项不同。看看任何一家大型科技公司:它的首要优先事项与AGI公司的首要优先事项不同。你的第一优先级决定了你的行动和结果,并最终定义了生态中的不同关系。

张晓军:为什么巨头分散投资多家基础模型公司,而不是重注一家?

杨植麟:这是阶段问题。未来会有更多整合——公司数量会减少。

张晓军:有人说基础模型公司的终局是被巨头收购。你同意吗?

杨植麟:我认为不一定。但它们很可能会有非常深入的合作。

张晓军:比如,如何合作?

杨植麟:OpenAI和微软是经典的合作模式。很多可以借鉴,有些可以改进。

张晓军:过去一年,创业的曲折在你身上体现在哪里?

杨植麟:有很多外部变量——资本、人才、GPU、产品、研发、技术。有高光时刻,也有需要克服的困难。以GPU为例。

反反复复:供应一度紧张,然后好转。最极端的时候价格每天都在变化——一台机器今天260,明天340,两天后又回落。这是一个不断变化的情况。你必须密切关注。价格不断变化,所以策略也不断变化:走哪个渠道,是买还是租——有很多不同选择。

张晓军:是什么推动了这些波动?

杨植麟:地缘政治原因;生产本身是分批的;市场情绪也起作用。我们观察到许多公司开始退还GPU,意识到他们不一定需要训练这个模型。随着市场情绪和人们决策的变化,供需也随之变化。好消息是最近整体供应大幅改善。我个人的判断是,至少未来一两年,GPU不会成为主要瓶颈。

张晓军:你似乎经常思考组织问题。你是如何组建团队的?

杨植麟:我们的招聘方式演变过。AGI人才在全球都极为有限,有相关经验的人很少。最早期的招聘画像专注于寻找有直接相关技能的天才。这被证明非常成功。那些以前“对模型动过手术”并有第一手训练超大规模模型经验的人,可以非常快速地构建东西。Kimi的推出包括在内——资本效率和组织效率实际上非常高。

张晓军:花了多少钱?

杨植麟:相当小的数字——与许多其他支出相比,是用小钱办大事。很长一段时间我们只有30-40人。现在80人。我们追求人才密度。

后来人才画像变了。早期我们招聘天才,认为他们的天花板高——公司的天花板由员工的天花板决定。后来我们补充了更多维度的团队成员——产品运营方面的人、领导型人才、能把事情做到极致的人。现在是一个更完整、有韧性的能战斗的团队。

张晓军:经过一年中国基础模型创业,你如何评估迄今为止的里程碑成果?

杨植麟:我们已经造了一个火箭原型,现在正在试射。我们组建了团队,搞明白了一些燃料配方,差不多能看到一个初具雏形的PMF。可以说我们已经迈出了登月的第一步。

§ 36

Zhang Xiaojun: What do you think of Yann LeCun’s position? He’s not optimistic about the current technical route—he believes self-supervised language models cannot acquire true knowledge of the world, and that as models scale, the probability of errors—machine hallucinations—will only grow. He has proposed the idea of a “world model.”

Yang Zhilin: There’s no fundamental bottleneck. When the token space is large enough, it becomes a new kind of computer that solves general problems—and that is a general world model.

An important point behind his statement: everyone can see the current limitations. But the solution doesn’t necessarily require an entirely new framework. The only thing that works in AI is next-token prediction plus the scaling law. As long as the tokens are complete enough, everything is doable. Of course, the problems he points out do exist today—but you solve them by making the token space very general. That’s all.

Zhang Xiaojun: So he’s magnifying the limitations.

Yang Zhilin: I think so. The underlying first principle is sound—it’s just that some small technical problems remain unsolved.

Zhang Xiaojun: What do you think of Geoffrey Hinton, the godfather of deep learning, repeatedly calling attention to AI safety?

Yang Zhilin: His focus on safety actually shows he has enormous confidence in the coming improvement of technical capabilities. The two are opposites.

Zhang Xiaojun: How do you solve the hallucination problem?

Yang Zhilin: Still the scaling law—it’s just that what you scale is something different.

Zhang Xiaojun: How likely is it that, in the end, the scaling law turns out to be a dead end?

Yang Zhilin: The probability is approximately zero.

张晓军:你怎么看Yann LeCun的立场?他对当前技术路线不乐观——他认为自监督语言模型无法获得真正的世界知识,随着模型扩展,错误概率——机器幻觉——只会增加。他提出了“世界模型”的想法。

杨植麟:没有根本瓶颈。当token空间足够大时,它就成为了一种解决通用问题的新计算机——那就是通用世界模型。

他表述背后的一个重要点:每个人都能看到当前的局限性。但解决方案不一定需要全新的框架。AI中唯一有效的是下一个token预测加规模法则。只要token足够完整,一切皆可行。当然,他指出的问题今天确实存在——但你可以通过使token空间非常通用来解决。仅此而已。

张晓军:所以他夸大了局限性。

杨植麟:我认为是的。底层第一性原理是成立的——只是一些小的技术问题尚未解决。

张晓军:你怎么看深度学习之父Geoffrey Hinton反复关注AI安全?

杨植麟:他对安全的关注实际上显示他对即将到来的技术能力提升有极大信心。两者是对立的。

张晓军:你如何解决幻觉问题?

杨植麟:仍然是规模法则——只是你扩展的东西不同。

张晓军:最终规模法则变成死胡同的可能性有多大?

杨植麟:概率大约为零。

§ 37

Zhang Xiaojun: What do you think of the view of your CMU alumnus Qi Lu: OpenAI will definitely be bigger than Google—it’s just a question of whether by one, five, or ten times?

Yang Zhilin: The most successful AGI company of the future will definitely be bigger than every company today—of that there’s no doubt. Ultimately, it could be a matter of double or triple GDP. It may not be OpenAI; it could be another company. But there will certainly be such a company.

Zhang Xiaojun: If you happened to become the CEO of this AI empire, what would you do to protect humanity?

Yang Zhilin: Thinking about that question now still lacks some preconditions. But we would certainly be willing to cooperate with and learn from different actors in society, including putting more safety measures into the models.

张晓军:你怎么看你的CMU校友陆奇的观点:OpenAI一定会比谷歌大——只是大1倍、5倍还是10倍的问题?

杨植麟:未来最成功的AGI公司一定会比今天所有的公司都大——这毫无疑问。最终,可能相当于两倍或三倍的GDP。不一定是OpenAI;可能是另一家公司。但肯定会有这样一家公司。

张晓军:如果你恰好成为了这个AI帝国的CEO,你会如何保护人类?

杨植麟:现在思考这个问题还缺少一些前提条件。但我们肯定会愿意与社会中的不同行为者合作和学习,包括在模型中增加更多安全措施。

§ 38

Zhang Xiaojun: What are your goals for 2024?

Yang Zhilin: First, technical breakthroughs—we should now be able to build a model far better than in 2023. Second, users and product—I hope for more users at scale and stronger retention.

Zhang Xiaojun: What are your predictions for the global foundation-model industry in 2024?

Yang Zhilin: More capabilities will appear this year, but the landscape won’t look much different from today—the top few players will still lead. On capability, there should be some fairly big breakthroughs in the second half of the year, many coming from OpenAI; it definitely has a next-generation model—maybe 4.5, maybe 5. That feels like a high-probability event. Video generation models can definitely keep scaling.

Zhang Xiaojun: And your predictions for China’s foundation-model industry in 2024?

Yang Zhilin: First, you’ll see new, distinctive capabilities emerge. Chinese models—because of earlier investment and having the right teams—will achieve world-leading capabilities in certain dimensions. Second, products with much larger user bases will appear—that’s highly probable. Third, there will be further consolidation and divergence in route choices.

张晓军:你2024年的目标是什么?

杨植麟:首先,技术突破——我们现在应该能够构建一个比2023年好得多的模型。其次,用户和产品——我希望有更大规模的用户和更强的留存。

张晓军:你对2024年全球基础模型行业有什么预测?

杨植麟:今年会出现更多能力,但格局不会与今天有太大不同——领先的少数玩家仍将领先。在能力上,下半年应该会有一些相当重大的突破,很多来自OpenAI;它肯定有下一代模型——可能是4.5,可能是5。这感觉是高概率事件。视频生成模型肯定可以继续扩展。

张晓军:你对2024年中国基础模型行业的预测呢?

杨植麟:首先,你会看到新的特色能力出现。中国模型——因为较早的投资和合适的团队——将在某些维度达到世界领先水平。其次,会出现用户基数大得多的产品——这概率很高。第三,进一步整合和路线选择的分化。

§ 39

Zhang Xiaojun: In starting this company, what’s the one thing you fear most?

Yang Zhilin: Nothing much, really—you just have to charge forward fearlessly.

Zhang Xiaojun: Anything you’d like to say to your peers?

Yang Zhilin: Let’s keep at it together.

Zhang Xiaojun: Name one question about the foundation-model industry that you don’t yet know the answer to but most want to know.

Yang Zhilin: I don’t know what the ceiling of AGI looks like—what kind of company it will produce, and what kind of products that company will create. That’s what I most want to know right now.

Zhang Xiaojun: As AGI keeps developing like this, what’s the one thing you’d least want to see?

Yang Zhilin: I’m fairly optimistic about it. It can take human civilization to the next stage.

Zhang Xiaojun: Has anyone ever said you’re too much of an idealist?

Yang Zhilin: We’re very down-to-earth, too. We’ve actually built some real things—we’re not just talking.

Zhang Xiaojun: If the money you’ve raised today were the last money you’d ever get, how would you spend it?

Yang Zhilin: I hope that never happens, because we’ll need a lot more money in the future.

Zhang Xiaojun: If you don’t make it, would you consider yourself a failure?

Yang Zhilin: It wouldn’t matter that much—I accept that failure is a possibility.

This endeavor has already completely changed my life, and I am full of gratitude.

张晓军:创办公司,你最害怕什么?

杨植麟:没什么可怕的——你只需要无畏地向前冲。

张晓军:有什么想对同行说的吗?

杨植麟:一起坚持。

张晓军:说出一个你还不知道答案但最想知道的基础模型行业问题。

杨植麟:我不知道AGI的天花板是什么样的——会产生什么样的公司,那家公司会创造什么样的产品。这是我现在最想知道的。

张晓军:随着AGI这样发展,你最不希望看到什么?

杨植麟:我对此相当乐观。它可以把人类文明带到下一阶段。

张晓军:有没有人说过你太理想主义?

杨植麟:我们也很务实。我们确实做出了一些实际的东西——不只是说说。

张晓军:如果今天融到的钱是你最后能拿到的钱,你会怎么花?

杨植麟:希望这永远不会发生,因为未来我们还需要更多钱。

张晓军:如果没做成,你会认为自己失败吗?

杨植麟:没那么重要——我接受失败是可能的。

这次尝试已经彻底改变了我的人生,我充满感激。

Open source ↗