2026-07-30

He Gave Away 2.8 Trillion Parameters, Then Left One Gate in the License — Yang Zhilin, and Why He Still Isn't on This List

On July 27, Moonshot AI uploaded the complete weights of Kimi K3 to Hugging Face.

1.42 TiB, 96 shards. 2.8 trillion total parameters, 104 billion active parameters, a 1-million-token context window. Of its 93 layers, 69 run Moonshot’s own Kimi Delta Attention and the other 24 use Gated MLA; there are 896 routed experts, 16 picked per token, plus 2 shared experts.

This is the largest open-weight model in the world right now. You can carry the whole thing home.

K3 went live as a service on July 16. Eleven days later the weights were given away. In those eleven days it scored 93.5 on GPQA Diamond, 67.5 on DeepSWE, 91.2 on BrowseComp, and 94.5 on MCPMark-Verified.

The parameters and the benchmark numbers have been passed around all day. Three other things sitting in the same materials have barely been mentioned, and all three are product decisions.

1. The real design isn’t in the model, it’s in the license

Everyone is saying “Kimi K3 went open source.” It is not MIT.

Moonshot wrote its own Kimi K3 License. The body of it does read like MIT — use it, copy it, modify it, distribute it, sublicense it, sell it, deploy it, fine-tune it, build derivative models on it.

Then come two exceptions.

One: selling it as a service means coming back to the table. If you turn it into model-as-a-service (MaaS) — giving third parties substantial control over inputs, parameters, or training data — and your company clears $20 million in revenue over any consecutive 12 months, you have to sign a separate agreement with Moonshot.

Two: get big and you carry the name. A commercial product with more than 100 million monthly active users, or more than $20 million in monthly revenue, has to display “Kimi K3” prominently in the interface.

The exemptions: purely internal use isn’t restricted, and going through Moonshot’s own products or certified inference partners isn’t restricted either.

Read those three passages together and the shape of the license comes out:

Small players use it freely. The day you’re making real money off it, either come back and split it, or run my billboard.

In my years doing product work, the open source strategies I’ve seen come in two flavors: genuinely wide open (and then you watch the cloud vendors monetize it while you get nothing), or a crippled edition released for show (and then nobody wants to build on it).

This license is a third road, and the incision is cut precisely. The gate sits at the point where you have already made money — before that line, friction is zero, and individual developers, startups, and researchers help themselves. Anyone who crosses it had the money to pay in the first place.

Look again at that second clause, the one requiring “Kimi K3” to be shown prominently. It doesn’t collect cash. It collects something else — it turns every large product running on K3 into a Moonshot billboard. Open source buys ecosystem, ecosystem buys brand, brand buys pricing power, and that chain has been written into the license as legal text.

There’s a move here you can copy directly: when you’re trading “free” for scale, work out in advance the exact moment it stops being free, and write that moment down as a line the other side can check for themselves. A vague free tier always ends in a mess of arguing.

The sequence was designed too. K3 shipped as a hosted service on July 16; the weights came 11 days later. In those 11 days the benchmarks got run, the word of mouth built, daily revenue went up sixfold — by the time the weights went out, it wasn’t a model waiting to be validated, it was a model that had been proven.

Run it the other way around: weights and service on the same day, everybody self-hosts immediately, and you get neither those 11 days of revenue nor that first batch of real usage data. Open source isn’t the same as not wanting money. It’s collecting the money and the data first, then putting the code out.

2. They just took first place, and used the release materials to say they’re behind

This one is rarer.

In their own release materials, Moonshot concedes that K3’s overall performance still trails Claude Fable 5 and GPT-5.6 Sol, and notes that there is a gap in user experience.

Not only that. They also volunteered a second point: different models were evaluated using different agent harnesses — meaning the benchmark environments are not fully comparable and the conclusions carry uncertainty.

They had just built the largest open-weight model on earth and topped several hard evals, with the whole world watching, and in their own announcement they said two things that work against them: the experience still isn’t as good as the competition, and the way these scores were compared is itself soft.

Tesla’s Q2 call yesterday was the mirror image. Musk said the humanoid robot demos going viral online right now are “mostly teleoperated or run off a pre-arranged script,” that a robot capable of genuinely performing general tasks on its own hasn’t appeared yet, and that Optimus will be the first — while reporting compiled that same week showed that at the 2024 Warner Bros. event, the Optimus units pouring drinks were being operated by engineers in motion-capture suits and VR headsets, and that at the Palo Alto headquarters a robot that falls over still needs a hoist to get it back on its feet.

Same week. One says everyone else’s is remote control and scripts. The other says our experience still isn’t as good as the competition’s.

Both plays have their logic. Musk’s expectation management has bought Tesla real time paid for in real money. Moonshot’s candor may simply be because it faces developers, and a developer can run the truth for himself inside a day — in front of that audience, bragging costs more than it returns.

But for a product manager, what you take away is the same either way: every sentence you say gets checked by users against today’s product. The only variable is how fast your users can check it. The faster they can, the better honesty pays.

And developers are the fastest-checking users there are.

3. Three times cheaper does not mean you spend a third as much

The price is the part that got passed around the most, and it’s the easiest to read wrong.

K3’s international pricing is $3 per million input tokens and $15 per million output tokens. Against Claude Fable 5’s $10 input and $50 output, that’s more than three times cheaper on paper. In the China region it’s ¥2 for cached input, ¥20 uncached, and ¥100 for output.

Then there’s the detail that got buried in the press coverage:

Multiple testers report that K3 burns more tokens than Fable to finish the same task.

That one line puts a discount on the “three times cheaper” above it. Unit price is not cost. What you pay is the total spend to get one thing done, which equals unit price times the number of tokens it takes to get it done.

This is a textbook product trap, and it isn’t confined to AI: you optimized a number the user can see, and the price of it is hidden in a number the user can’t compute.

I’ve walked into the same hole myself. Building my publishing pipeline, I was optimizing “time to publish on a single platform,” squeezing every step shorter, and the whole thing got slower — because the shortened steps tripped the platforms’ anti-automation, and the retry count went up. The local metric looked better and end-to-end got worse.

So if you’re evaluating a switch to K3, don’t read the price list on the website. Run your own real workload through it and total up the cost of one completed job. That number may still favor K3 — the discount is large enough to survive a lot — but it should be a number you computed, not a number they printed.

The demand is real: after K3 launched, Moonshot’s daily revenue rose at least sixfold; June ARR hit $300 million, up from $200 million in April.

4. Who he is

Yang Zhilin, born 1992 in Shantou, Guangdong, top science scorer in Shantou’s college entrance exam. Undergrad at Tsinghua, transferred into the Yao Class in his second year, graduated first in his year in 2015. Then a PhD at Carnegie Mellon under Ruslan Salakhutdinov, Apple’s first director of AI, and Google’s William W. Cohen.

He is first author on both Transformer-XL and XLNet. XLNet beat BERT on 20 tasks at the time and set best results on 18 of them. He is among the most cited NLP researchers under 35 in China.

One more thing that has nothing to do with the technology and that I think matters: he fronted a rock band at Tsinghua called Splay, as lead singer and lyricist.

In March 2023 he founded Moonshot AI with Zhou Xinyu and Wu Yuxin. In February 2024 Alibaba led a $1 billion round; in August, Tencent and Gaorong put in another $300 million.

The gate in the license, the candor in the announcement, the trade-offs in pricing — none of those three are researcher thinking. A researcher’s default move is to max the score and publish the paper. Setting a precise commercial gate inside a license, and volunteering that you’re behind the competition at your most triumphant moment, are product and business judgments.

A first author on XLNet doing both of those things at once is a lot more complicated than the label “technical guy with strong papers.”

5. So why isn’t he on this list

I maintain a list called The 100 Product Managers Who Changed the World, scored across six dimensions — vision, insight, taste, business, scale, originality — weighted into an OVR. Yang Zhilin isn’t on it.

The rule works like this: the list has only 99 entries, and slot 100 is deliberately left blank, reserved for the reader. To put someone in, you have to take someone else out.

Bumping a person off the list to chase today’s news would produce a list that tracks the news cycle instead of product history. So this piece doesn’t add him.

Not adding him isn’t the same as saying he doesn’t deserve it, and the standard should be stated out loud.

The person already on the list in the same bracket is Liang Wenfeng, OVR 91 (vision 95 / insight 84 / taste 86 / business 88 / scale 92 / originality 95). DeepSeek was an event that changed the cost structure of the entire industry — it broke the consensus that frontier models have to be expensive, and pricing and open source strategy moved worldwide in response.

Holding the same ruler up to Yang Zhilin, my read is: vision and originality are already enough to make the list; scale and business are still short of breath.

Put plainly: the most valuable part of his score is happening, not delivered.

That’s the difference between a ranking and the news. The news records what is happening; a ranking records what has settled. Putting him in today would be me betting on next year’s result, not stating a fact.

But this list is alive. Allen Zhang and Steve Jobs probably aren’t moving, while the middle and the back of it change constantly — when Linus was added back in, Will Wright came out. I plan to re-rank the entire list in the second half of the year, and by then what the K3 open-weight ecosystem actually grew, how much that commercial gate actually collected, and whether Moonshot survived the White House will all be plain facts.

If those have all been delivered by then, putting him on the list isn’t a bet, it’s a record.

The White House part is still hanging. After the K3 weights went out, White House CTO Michael Kratsios publicly accused Moonshot of distilling Anthropic’s Fable 5, saying a distinction has to be drawn between normal AI distillation and “large-scale, covert, industrial-scale distillation designed to steal US proprietary technology”; Treasury Secretary Scott Bessent said the government “has the ability to impose sanctions.”

An open-weight model that can draw out a country’s CTO and its treasury secretary at the same time already carries weight. But how that line plays out directly determines what his company’s scale dimension is worth next year.

Closing

“Free” was never the decision. “From what moment does it stop being free” is the decision. Most people doing open source, freemium, or a trial period only work out the first half of that sentence and leave the second half fuzzy — and then either they can’t collect when it’s time to collect, or they throw up a paywall at the wrong moment and scare the ecosystem off.

Yang Zhilin nailed that line down inside the license: $20 million in revenue, 100 million monthly actives. Whoever crosses it knows they crossed it. No negotiation, no arguing.

I didn’t think this carefully when I built SoloMD. It’s free right now, MIT licensed, 551 stars, 21 external PRs — the ecosystem exists, but if someone turns it into a business tomorrow, I don’t have a single line I can point them to.

(All scores and rankings in this piece were produced by Claude (AI); the method is explained on the rankings page. The assessment of Yang Zhilin is this article’s own analysis and is not counted in the list.)

Discussion

No login needed. Be kind.
Loading…