Glean 拾遗
日刊 /2026-08-13 / Grok 4.6 与 DeepSeek V4 Pro 同日实测:准 Fable 级能力,价格打到底

Grok 4.6 与 DeepSeek V4 Pro 同日实测:准 Fable 级能力,价格打到底

原文 x.com 收录 2026-08-13 10:47 阅读 6 min
AI 解读

作者以个人开发者视角,记录 DeepSeek V4 Pro 与 Grok 4.6 同夜发布后的选型变化。关键数据:DeepSeek V4 Pro 百万输出 Token 0.87 美元,约为 Claude Fable 5 的 1/57,Terminal Bench 2.1 得分 87.9(Fable 5 为 88.0),DeepSWE 从预览版 12.8 升至 62.7;Grok 4.6 输入/输出定价 2/6 美元每百万 Token,作者实测连续开发 4 小时仅消耗 SuperGrok Heavy 周额度 2%。作者同时承认不可能三角未完全打破:DeepSeek 仍缺多模态且速度一般,Grok 在纯开发与 Agent 场景逊于 Fable 5。适合关注大模型选型、成本优化与开发速度的 AI 工程师。

原文 6 分钟
原文 x.com ↗
§ 1

Overnight, DeepSeek V4 Pro and Grok 4.6 push the world's LLMs onto a new kill line.

Today can only be described as a magical day.

Two things happened:

One was DeepSeek V4 Pro's official release.

The other was Grok 4.6 officially launching.

Both happened overnight, within a two-hour window.

And the amusing part is that both are models of similar parameter scale — one at 1.6T, the other at 1.5T — yet both, tonight, edged close to the Claude Fable 5 experience.

一夜之间,DeepSeek V4 Pro和Grok 4.6把全球大模型推入了新的斩杀线。

今天,可以说是神奇的一天。

两个事:

一个是DeepSeek V4 Pro正式版发布。

另一个是Grok 4.6正式发布。

这是一夜,2个小时之内发生的事。

而且有趣的是,这两全都是同等参数模型。

一个1.6T,一个1.5T,但是都在今晚,逼近了Claude Fable 5的体验。

§ 2

Then, after trying them out tonight, I voted with my wallet.

For the first time in my life, I dropped $300 on a SuperGrok Heavy membership for Grok.

At the same time, I turned around and topped up the DeepSeek API with 1,000 yuan (~$140) of credit.

A little something for each side.

然后呢,我今晚在体验完以后,也用脚投票了。

我人生第一次为Grok怒氪了300美元的SuperGrok Heavy会员。

同时,调头给DeepSeek的API充了1000块钱的额度。

雨露均沾了一下。

§ 3

And with these two releases, I think the LLM market is about to get a razor-sharp kill line.

This round will likely eliminate a great many players.

It's been barely three or four months since Claude was so far ahead that the gap felt hopeless; now leadership changes hands every few weeks.

So, at this crazy moment, I'm also updating my everyday AI lineup:

First: for my daily main Vibe Coding work, I'm moving to Grok Build + Grok 4.6 — I'll explain why later.

Second: ChatGPT and Codex.

They'll keep handling the hardest, absolutely-cannot-fail mega-tasks, plus all my daily office work, creative and content production, and browser/computer control.

Third: DeepSeek V4 Pro (stable).

Because of its unbeatable value, I'll hand it a large volume of automation and async tasks that don't need real-time feedback, plus some automation Agent work inside products I build, like AIHOT.

That's the combination I've settled on after tonight's reshuffle.

Who knows whether it'll change once new models land — but for now, this genuinely is the setup that feels best to me.

而且,这两个模型发布以后,我觉得现在这个大模型的市场,会有一层非常鲜明的斩杀线了。

这一轮大模型,可能又会被淘汰掉很多很多玩家。

从曾经的Claude遥遥领先,让人感受到绝望的差距,到现在,各领风骚好几周,也才不过短短3、4个月而已。

所以在今天这个疯狂的时间点,我也更新一下我自己现在的常用AI组合:

首先,我未来的日常主力Vibe Coding产品,会使用Grok Build + Grok 4.6了,理由我后面会说。

第二,ChatGPT和Codex。

它们会继续负责最困难、绝对不能出任何差池的超大型任务,还有日常的所有办公任务、创意、内容生产,以及浏览器操控和电脑操控相关。

第三,DeepSeek V4 Pro正式版。

因为无与伦比的性价比,我会把大量不要求实时反馈的自动化任务、异步任务,还有我开发的产品比如AIHOT里面的一些自动化Agent工作交给它。

这就是我今晚重新分配完以后,暂时形成的模型组合。

鬼知道后面的新模型诞生后,会不会变,但是目前来说,这确实就是我自己用着最舒服的最优解。

§ 4

The reason behind this split is simple.

I've always believed LLMs face an impossible triangle — three things that are brutally hard to satisfy at once:

Performance.

Price.

Speed.

This is the triangle that, historically, almost nobody could fill.

People talk about the first two all the time, but the last one — speed — is constantly forgotten.

Because with speed, you can only truly feel it if you're using these models every day. It depends not just on the model, but also on the vendor's inference capability and compute reserves.

If, like me, you spend close to ten hours a day with various Agents working nonstop, speed is sometimes the single most decisive factor.

之所以这么分配,理由特别简单。

因为我一直觉得,大模型其实存在一个很难同时满足的不可能三角。

性能。

价格。

速度。

这就是过去几乎满足不了的不可能三角。

前面两项大家可能讨论的最多的,但是最后一条常常被人遗忘。

因为速度这事,你真的只有天天用,你才能感受的到,因为不止跟模型有关系,也跟厂商的推理能力、算力储备有关系。

如果你和我一样,每天接近十个小时的时候,几乎都在让各种Agent不停歇的干活,速度有时候就是一个至关重要的因素。

§ 5

I can barely tolerate the speed of Codex in normal mode with GPT-5.6 Sol anymore.

A task taking ten-odd minutes is common; half an hour as a starting point isn't rare; some tasks routinely stretch to a full hour.

What I hate most is getting feedback every ten or thirty minutes — it forcibly chops my time into little fragments, and while I wait, bored out of my mind, all I can do is scroll X...

Which is why my X-scrolling time has exploded lately...

So lately I've been forced to keep Codex in Fast mode around the clock.

And the way my quota burns — I'm on the $200 membership, and after a couple of casual days, the quota collapses.

This is a very, very, very real situation.

It's the impossible triangle I've been facing daily for a long time.

我现在其实已经快接受不了Codex用GPT-5.6 Sol正常模式的速度了。

一个任务十几分钟很常见,半小时起步也不稀奇,有些任务常年能干到一个小时。

我最讨厌的就是十几分钟或者半小时一个反馈,硬生生的把我的时间切成一块一块的碎片,然后这个时间我又百无聊赖,只能在那刷X。。。

导致我最近刷X的时间都暴涨。。。

所以最近我用Codex,常年被迫开着Fast模式。

然后我的额度那烧的,我是200刀会员,随便跑两天,额度就撑不住了。

这是一个非常非常非常现实的情况。

也是我过去很长时间里,我每天真实面对的模型的不可能三角。

§ 6

The smart ones — Claude Fable 5, GPT-5.6 Sol — are usually expensive and slow.

The cheap ones are often fast, but their performance falls short; DeepSeek V4 Flash, for example, is fine for everyday use, but a mediocre coder like me still needs stronger model capability to raise my ceiling.

Fast and powerful? That's basically what the various Fast modes sell you, and they're anything but cheap.

This is the impossible triangle of LLMs — and it's also the industry's kill line.

聪明的,比如Claude Fable 5、GPT-5.6 Sol,经常是又贵又慢。

便宜的,速度确实经常也快,但是性能又经常差一截,比如DeepSeek V4 Flash,日常确实够用,但是对于我这种菜逼来说,我还是需要更强的模型能力来拔高我的上限。

速度快性能又好,那基本就是各种Fast模式,你也便宜不了。

这就是大模型的不可能三角,更是整个行业的斩杀线。

§ 7

But in the early hours of today, DeepSeek V4 Pro and Grok 4.6 ripped a hole in this impossible triangle — from two different directions.

At current pricing, DeepSeek has squeezed Fable-class Agent capability down to $0.87 per million output tokens. They did announce a price increase ahead of time, and I don't know what it'll settle at, but I doubt it'll exceed $2.

Grok 4.6, meanwhile, turned near-Fable-5 comprehensive capability into a $2 input / $6 output price — and handed me a real-world development experience that's absurdly fast.

但是,就在今天凌晨,DeepSeek V4 Pro和Grok 4.6从两个方向,一起把这个不可能三角,逐渐的撕开了一个口子。

按照目前定价,DeepSeek把Fable级Agent能力压到了百万输出Token 0.87美元,虽然之前发过涨价预告,不知道涨价完以后多少,但是我觉得应该超不过2美元。

而Grok 4.6更是把接近Fable 5的综合能力变成了2美元输入、6美元的输出价格,而且还给了我一个快到有点离谱的真实开发体验。

§ 8

So a new kill line has formed. I've added this round's DeepSeek V4 Pro and Grok 4.6 to that famous LLM kill-line chart. Since DeepSeek V4 Pro's score hasn't been published yet, I estimated from the benchmarks and figured it lands around 56–58.

The LLM world is left with no survivors — total coverage.

于是,新的斩杀线形成了,我把那个著名的大模型斩杀线的图补了这次的DeepSeek V4 Pro和Grok 4.6,DeepSeek V4 Pro的分因为还没出,我大概根据跑分,算了一下,估计就是56~58左右。

大模型世界,再无活口,全面覆盖。

§ 9

First, let's look at DeepSeek V4 Pro's benchmark scores.

The numbers look dramatic, but out of all those scores, you only need to watch two.

The first is Terminal Bench 2.1. In simple terms, it makes the model enter a computer terminal by itself — installing the environment, typing commands, handling errors — and actually completing the task.

DeepSeek V4 Pro's preview version scored just 72.1; the official release jumped straight to 87.9, while Fable 5 sits at 88.0 — a gap of 0.1.

The second is DeepSWE.

This is the benchmark I focus on most. It's much closer to real software development: the model has to read and understand a code project on its own, then solve complex engineering problems inside it.

DeepSeek rocketed from 12.8 to 62.7 — a jump of nearly 50 points in one go.

Fable 5 scores 70, so DeepSeek still trails. But it's gone from being almost useless to dramatically stronger — it can actually get work done now.

先看DeepSeek V4 Pro的跑分。

跑分看着是比较夸张的,这么多跑分,其实你就关注两个就行。

一个是Terminal Bench 2.1,你可以把它简单理解成,让模型自己进入电脑终端,装环境、敲命令、处理报错,最后真的把任务做完。

DeepSeek V4 Pro预览版只有72.1,正式版直接涨到了87.9,而Fable 5是88.0,只差0.1。

第二个,是DeepSWE。

这个是我重点看的一个评测集,它需要更接近真实的软件开发,模型需要自己读懂一个代码项目,再处理里面复杂的工程问题。

DeepSeek从12.8一路冲到了62.7,一次涨了将近50分。

Fable 5是70分,所以DeepSeek依然有差距,可它已经从几乎干不了,还是大幅强化了很多,能干活一点了。

§ 10

The same trend shows up across codebase understanding, cybersecurity, automation, and tool calling: once it's given tools and actually allowed to operate in an environment, DeepSeek V4 Pro quickly closes in on Fable 5 — and even beats it on a few specific projects.

But its weaknesses are just as clear.

When it can't use tools and has to rely purely on the model's own knowledge and reasoning, it still trails Fable 5 by a dozen-plus points. On the hardest software engineering and complex full-stack tasks, it also lags by a few points.

So I'd rather call it a near-Fable-class model for Agent and coding work.

Its combat power inside tool environments is now first-tier, but its pure knowledge, hard software engineering, and long-horizon stability all felt a bit shaky in my own testing.

There's also one fatal point: it still doesn't do multimodal.

But that so-called fatal flaw doesn't matter much at this insane price and cache-hit rate.

其他什么代码仓库理解、网络安全、自动化和工具调用,趋势也都差不多。

只要给它工具,让它真的开始操作环境,DeepSeek V4 Pro就会迅速逼近Fable 5,个别项目甚至还能超过。

但是它的短板也很清楚。

不使用工具,只考模型自身知识和推理时,它和Fable 5还会差十几分。碰到最困难的软件工程和复杂全栈任务,也依然会落后几分。

所以,我更愿意把它叫作Agent和Coding领域的准Fable级模型。

它在工具环境里的战斗力已经进了第一梯队,但是纯知识、困难软件工程、和超长期稳定性其实在我自己测试下来,也是感觉有点问题的。

特别还有个致命的点,就是还是没有多模态。

但是这个所谓的致命点,在如此离谱的价格和缓存命中率下,也不算啥了。

§ 11

Fable 5's API pricing is $10 per million input tokens and $50 per million output tokens.

DeepSeek V4 Pro currently charges $0.43 per million input tokens and $0.87 per million output tokens, with cache-hit input at just $0.0036.

That makes regular input roughly 23x cheaper.

Output roughly 57x cheaper.

And cache-hit input nearly 276x cheaper. Credit where it's due: DeepSeek's cache technology is straight-up black magic — not just cheap, but with an extremely high hit rate.

At this price, who the hell is going to complain, right? Of course, that assumes DeepSeek V4 Pro doesn't raise prices — and if the increase is modest, then paired with WorkBuddy in the future, I think that's the new god-tier combo.

Fable 5的API价格,是百万输入Token 10美元,百万输出Token 50美元。

DeepSeek V4 Pro当前是百万输入Token 0.43美元,百万输出Token 0.87美元,缓存命中输入只有0.0036美元。

普通输入大约便宜23倍。

输出大约便宜57倍。

缓存命中输入的差距接近276倍,有一说一,DeepSeek的缓存技术实在是过于黑科技了,不仅便宜,命中率还极高。

这个价格,还挑个屁对吧,当然,前提是DeepSeek V4 pro不涨价,如果涨价不多的话,那未来搭配WorkBuddy,我觉得就是新一代的真神组合了。

§ 12

I plugged DeepSeek V4 Pro straight into Claude Code and ran a bunch of real development and Agent tasks.

The capability is genuinely good — especially development and tool calling. I can clearly feel it's in a completely different state from the preview version, which I complained about quite a bit. Now at least it's usable; unless you're building something huge, you basically can't tell the difference.

Economically, this model is the new-generation kill line by itself.

我把DeepSeek V4 Pro直接接进Claude Code,实际跑了一些开发和Agent任务。

能力确实是不错,尤其是开发和工具调用,我能很明显感觉到它和预览版已经完全是两种状态了,预览版我其实吐槽过很多,现在至少是可以的,不开发大活你基本也分辨不出来啥区别。

这个模型,在经济价值上,直接就是新一代斩杀线了。

§ 13

But here's the catch — the impossible triangle still holds: DeepSeek V4 Pro's actual development speed is really quite average.

Especially when I'm using Grok 4.6 at the same time: over there, three tasks are already done; over here, one hasn't even finished. The contrast in speed is stark.

However, if you don't need real-time responsiveness and speed, or — like me — you're just running async automation, then I think this is a perfect model for today.

In the past, I'd habitually offload all those automation tasks to Codex.

But today, I've basically switched all of them to DeepSeek V4 Pro.

但是呢,问题就来了,那个不可能三角的定理还是在,就是DeepSeek V4 Pro实际开发速度,真的挺一般的。

尤其是当我同步在用Grok 4.6,那边都跑三任务了,这边一个还没跑完,这个速度的对比,还是挺强烈的。

不过如果你对任务的实时性和速度没有那么高的要求,或者跟我一样,就是跑异步自动化,那我觉得,这就是一个当今完美的模型。

以前我可能会习惯性把那些自动化任务任务全部挂在Codex上。

但是今天,我基本全切到DeepSeek V4 Pro了。

§ 14

And then there's my new favorite, Grok 4.6.

Honestly, it's the only model out there today that I feel has nearly maxed out the impossible triangle: comprehensive performance rivaling Claude Fable 5, blazing speed, and a price that's higher than a price-slasher like DeepSeek, but several times cheaper than today's OpenAI and Anthropic flagship models.

然后就是我的新宠,Grok 4.6。

讲道理,这是当今唯一一个,我觉得快把不可能三角拉满的,综合性能媲美Claude Fable 5,速度快到飞起,价格比DeepSeek这种价格屠夫要贵,但是相比现在OpenAI和Anthropic的旗舰模型,直接就砍了好几倍。

§ 15

Grok 4.6's API starts at $2 per million input tokens and $6 per million output tokens.

That's 5x cheaper than Fable 5 on input, and roughly 8.3x cheaper on output.

You might look at the price and performance and think it's about the same as Kimi K3 — but you can't do the math that way, because the Musk camp has subscriptions.

Today I signed up for $300 SuperGrok Heavy, and there's a promo right now: subscribing to SuperGrok Heavy gets you a Cursor Ultra membership worth $200... for free.

So for $300, you get the quota of two memberships.

Cursor alone gives you access to a pile of Claude models and so on — I won't even go into that. Just on the SuperGrok Heavy itself... do you know how absurd its quota is?

Grok 4.6的API起步价格,是百万输入Token 2美元、输出Token 6美元。

输入比Fable 5便宜5倍,输出便宜大约8.3倍。

虽然你看价格和性能跟Kimi K3差不多,但是这个价格,你不能这么算,因为,老马家是有订阅的。

我今天开了300刀的SuperGrok Heavy,而且现在还有活动,你订阅SuperGrok Heavy,还可以直接送你Cursor的价值200刀的Ultra会员。。。

也就是300刀,你能享受两个会员订阅的额度。

Cursor里面可以用一堆的Claude模型啥的,我就不提了,就单说这个SuperGrok Heavy,你们知道它的额度有多离谱吗。

§ 16

Do you know how absurd its quota is?

I subscribed just past 2 AM, then used their Grok Build and Coding straight through to 6 AM.

Four hours total — and the quota only dropped 2%. It went from 1% to 2%, and that tick only happened around 5:30. So for the first three and a half hours, I was developing nonstop, plus running N Agents on Heavy for deep research, and I'd used only 1% of my weekly quota. Even knowing this first week includes double usage, that's still absolutely insane.

At that burn rate, it's cheap to the point of absurdity — and I haven't even counted the Cursor quota.

你们知道它的额度有多离谱吗。

我从2点多订阅完,一直用他们的Grok Build然后Coding到了早上6点。

4个小时,而这个额度,一共只消耗了2%,从1跳到2%,还是5点半左右涨的,也就是前3个半小时,我一直开发,还加上用Heavy开了N个Agent去深度调研,我只用掉了周额度的1%,即使我知道这是第一周包含了双倍用量,但这依然也太离谱了。

你要按这个额度算,绝对是便宜的没边的级别,我甚至还没有算上Cursor的额度。

§ 17

Its overall performance is extremely strong, too. Compared with Kimi K3, apart from a slightly weaker aesthetic sense on frontend work, I don't really feel much difference.

And on real-world tasks — building financial models, PPTs, legal cases, and the like — Grok can even win a few. But its weak spot remains development and Agent work, where it does fall a bit short of Fable 5.

So that's why I call it a near-Fable-class model.

综合性能也极强,基本跟Kimi K3比,除了前端审美差一些,其他的我感觉是没啥区别的。

而且在真实世界的任务中,比如做财务模型、PPT、法律案件之类的,Grok还能赢一些,但是短板依然还是开发和Agent上,这块确实比Fable 5还要差一些。

所以我说的是准Fable级的模型了。

§ 18

And its development speed is comfortable beyond words.

I developed continuously with Grok 4.6 for about four hours and barely ran into a task that took over 20 minutes.

Some tasks finished in as little as 3 minutes — with no drop in quality.

This is a speed I almost never experienced with Codex.

In the past, a task would run ten-odd minutes, half an hour, or even a full hour, and I'd be worn numb — spending most of the time just waiting. Today, developing with Grok 4.6, I felt, for the first time in ages, a genuine sense of interaction between me and my Agent.

Trust me: in many cases, speed is what truly turns a model's intelligence into productivity.

This impossible triangle — I think Musk pulled it off today. Grok 4.6 pulled it off.

That's why I said at the start that from today, I'm switching all my main daily development to Grok Build + Grok 4.6. I can sum it up in one word: bliss.

而它的开发速度,更是舒适的无以复加。

我用Grok 4.6连续开发了大概四个小时,几乎没有碰到超过20分钟的任务。

有些任务甚至3分钟就结束了,而且不减质量。

这绝对是我用Codex时几乎从来没有体验过的速度。

过去一个任务跑十几分钟、半小时甚至一小时,我人都快被磨麻了,绝大多数的时间,都在等,而今天,我用Grok 4.6开发,居然久违的感受到了我跟Agent之间交互的感觉。

相信我,在很多时候,速度也会把模型的智力,真正变成生产力。

这个不可能三角,我觉得在今天,老马是做到了,Grok 4.6做到了。

所以我在开头说,我会从今天开始,把日常的主力开发,全部切换到Grok Build + Grok 4.6上,爽怎就一个字了得。

§ 19

And Musk has thrown down another wild claim:

Grok 4.7 will surpass every model.

This time, I believe him.

而且老马又放出了狂话:

Grok 4.7要超越所有模型。

这一次,我信老马。

§ 20

So back to the very beginning: why did I say DeepSeek V4 Pro and Grok 4.6 would push all the LLMs onto the kill line?

Because the first thing a kill line severs is a model's right to keep being arrogant.

In the old LLM market, everyone could find room to survive.

The most capable could charge a premium and take their time.

The weaker ones could win users with low prices.

The fast enough ones could specialize in simple tasks.

Because the impossible triangle existed, every model had a shortcoming, and users could only accept all kinds of compromises.

But today, the market's minimum bar is different.

那回到最开头,为什么我会说,DeepSeek V4 Pro和Grok 4.6,会把一众大模型推入了斩杀线?

因为斩杀线首先斩掉的,是一个模型继续傲慢下去的资格。

过去的大模型市场,大家都可以找到自己的生存空间。

能力最强的,可以贵一点、慢一点。

能力弱一些的,可以靠低价抢用户。

速度足够快的,也可以专门承接简单任务。

因为不可能三角存在,所以每个模型都有短板,用户也只能接受各种各样的妥协。

但今天,这个市场的最低标准变得不一样了。

§ 21

DeepSeek V4 Pro is telling everyone that near-Fable-class Agent capability can cost just $0.87 per million output tokens (and yes, I know it'll rise later, but I'd guess it settles around $2).

Grok 4.6 is telling everyone that a model with overall performance close to Fable 5 can simultaneously have blazing development speed and a subscription quota that drops only a few percent after hours of heavy use.

When both of these exist at once, the models caught in the middle are going to be in serious pain.

Grok 4.6, in particular, is basically the gatekeeper of Diamond rank in this game.

If you can't beat these two, what's the point of using you?

The entire kill line — pushed forward by DeepSeek V4 Pro and Grok 4.6 — has jumped a huge step ahead.

This is the best of times.

And also the worst of times.

DeepSeek V4 Pro告诉所有人,准Fable级的Agent能力,百万输出Token可以只要0.87美元(当然我知道后面还会涨,但是估计也就2美元)。

Grok 4.6则告诉所有人,一个综合性能接近Fable 5的模型,可以同时拥有极快的开发速度以及一个高强度使用几个小时才掉百分之几的订阅额度。

当这两个东西同时出现以后,夹在中间的模型就会非常难受。

特别是Grok 4.6,基本在游戏里,现在就属于钻石分段守门员。

你要是打不过这两,我用你的意义在哪里?

整条斩杀线,被DeepSeek V4 Pro和Grok 4.6。

向前推了一大截。

这是一个最好的时代。

但,也是一个最坏的时代。

打开原文 ↗