# doaipm — DO AI PM · full content for LLMs > Become a product manager in the AI era. Speak it, and AI builds it (言出法随). > This file contains the full text of doaipm's articles so AI search engines can > retrieve and cite them directly. Canonical site: https://doaipm.com/en/ > Method: https://doaipm.com/en/method/ · Skill: https://github.com/zhitongblog/doaipm-skill Core ideas: trust the AI (don't tool-hunt); not knowing how to code is an advantage; high-fidelity first (build a real runnable prototype, not wireframes); the five-phase workflow Discover → Define → Design → Develop → Validate; and a safety net (no real secrets in prototypes, the human presses irreversible buttons, ask when unsure). Maintained by zhitong (智通) · https://zhitong.uk --- # The PM's AI Transition Guide 01 | Which "AI Product Manager" Are You Actually Trying to Become? URL: https://doaipm.com/en/blog/which-transition/ Published: 2026-08-13 Tags: AI Transition, AI Product Manager, Product Management, Career, Tencent Recruiting, Zhilian, Tech Commentary Tencent opened its 2027 campus recruiting on August 11, aimed at students graduating between January 2026 and December 2027. The announcement lists a batch of AI-native roles: AI full-stack engineer, agent development engineer, AI application engineer, AI algorithm engineer, AI product manager, AIGC art creation. The same announcement changed something else. Tencent's long-standing hiring creed used to be "ambitious, eager to learn, gets things done." This time three more specific clauses follow it: **proactively find and define problems, evaluate and verify AI output, turn ability into visible results.** Not one of the three is about large-model technology. And they weren't written for the AI product manager role alone. Tencent opened five job families this round — technology, product, design, marketing, and corporate functions — and the standard applies to everyone across all five. That stopped me. Because every "transition guide" I've read over the past year or so is about something else entirely. ## "AI product manager" names two different things at once Search for how a PM should transition into AI and the results are remarkably uniform: a learning roadmap — prompt engineering → RAG → agents → LLMOps — a salary table, and a signup link at the bottom. I started down that road too. It took me a few months to notice that what I needed wasn't on the map. The trouble is that one job title gets used for two different things: | | "AI product manager" | "Product manager in the AI era" | |---|---|---| | What it is | A **job**: building large-model and agent products | A **way of working**: using AI to build anything | | What it takes | RAG, fine-tuning, eval systems, Coze / Dify | Stating requirements clearly, judging what's worth building, shipping it yourself | | Who gets in | A limited number of seats at a limited number of companies | Anyone who already has work in front of them | | Share of content online | Almost all of it | Very little | Tencent's three clauses describe the right-hand column. The courses being sold teach the left. The left column isn't fake. It's real, and it's growing. ## Zhilian's numbers only mean something when you read them together Zhilian Recruitment published its *2026 AI Industry Talent Development Report* on July 16, based on first-half platform data. Look at role-level growth alone and AI PM looks excellent: - Agent-related technical talent: **+244%** year over year - **AI product manager: +87.7%** year over year - Data annotation / AI trainer: **+30.3%** - AI engineer: **+19.3%** The number is still climbing. From the same organization, vice president Li Qiang put Q1 growth at +81%; the half-year figure is 87.7%. His data covered 404 million job seekers and 15.49 million companies on the platform. But the same report carries another figure that belongs right next to it: **the entire AI industry's job postings grew only 10.6% year over year, while the number of job seekers grew 10.5%.** > The AI product manager role is growing fast, but it's growing inside a pool that expanded only 10.6%. 87.7% is a growth rate, not a volume. Doubling a small base can still leave you with a small absolute number. And the role growing far faster — agent development at +244% — is an engineering role that product managers can't simply move into. Postings up 10.6%, applicants up 10.5%. Those two numbers sit almost on top of each other: the door is getting wider, and the line outside it is lengthening at the same pace. ## The fastest AI hiring growth isn't at large-model companies Two more figures from the same report, broken out by industry: - AI engineer postings in new energy: **+38.2%** year over year - AI engineer postings in aviation, aerospace, and shipbuilding: **+23.5%** Both are well above the industry-wide 10.6%. A meaningful share of the growth, in other words, isn't at companies building large models. It's at traditional industries putting AI into their own operations. The number of companies hiring rose 24.8% — faster than the 10.6% growth in postings — which means more companies each opened a few seats, rather than a handful of firms hiring at scale. That has a fairly direct bearing on which way to move. If AI work is spreading into new energy, manufacturing, and aviation, the person those companies need may not be a product manager who understands large models. More likely it's a product manager who understands that industry and can put AI to work in it. That second person already has the domain knowledge. What's missing is the way of working. ## How narrow is the door? 54 job descriptions give you a read Someone went through 54 AI PM job descriptions at leading companies. The recurring requirements cluster tightly: prompt engineering, RAG, function calling, data alignment, SFT / RLHF, automated evaluation systems, low-code platforms like Coze and Dify, plus paper-reading and the ability to prototype in code. What employers weigh most heavily is **whether you have real experience shipping AI into production**. The reason isn't hard to see: AI projects run into a pile of problems that have nothing to do with technology — you can't get the data, a department won't cooperate, users simply don't trust what the model returns. Only people who've done it know where those holes are. That same analysis turned up a counterintuitive finding: **no agent project experience doesn't mean no shot.** Employers will credit hands-on work on low-code platforms, and they'll credit a background as a heavy user. So the road isn't sealed off. It just asks you to produce something, rather than produce proof of attendance. The catch is that "something" is exactly what a learning roadmap cannot hand you. ## "Evaluate and verify AI output" has no job-title prerequisite Back to Tencent's three clauses. The middle one is the one I kept staring at: **evaluate and verify AI output.** What makes it unusual is that it asks nothing of your job title, your employer, or your industry. Whatever work is in front of you right now, the moment you start doing it with AI, this begins the same day. It's also harder than it looks. AI will describe unfinished work as finished — not out of dishonesty, but because it doesn't know which parts it left undone either. Either you can spot it yourself, or you wait for users to tell you after launch. Last year I started trying to take things all the way to shippable on my own. Two of them have commit histories you can check: an ID-photo tool, 21 commits, June 13 through June 30, now in the App Store; and a PDF tool, 43 commits, the first on July 5 and the last on July 11, shipped for both iOS and macOS. I'm not citing those numbers to suggest it was hard. The opposite — most of the hard parts were absorbed by AI. The time went somewhere else: judging whether what it produced was usable, and finding the places I had to take over by hand. Tencent's phrase for this is "turn ability into visible results." *Visible* is the load-bearing word. "Familiar with large models" on a résumé can't be verified by the person reading it. A link they can click answers the question in thirty seconds. ## I can't tell you which road to take If your company happens to be building large-model products, or you want to work somewhere that does, the job route is real. The 87.7% is real. The job descriptions spell out what they want, and you can go fill those gaps. If the work in front of you has no direct connection to large models — and per that report, a lot of the growth is flowing toward exactly this kind of work — then grinding through RAG and SFT may pay less than actually building one idea all the way through, once. It's also possible these two roads were never meant to be read separately. On the same day Tencent listed "AI product manager" in its job table, it wrote "evaluate and verify AI output" into the standard for all five job families. Same announcement, same day. Here's what I still haven't worked out: if nobody at your company asks you to produce this kind of thing, and your performance review doesn't look at it, what keeps you going? I've gotten this far by building my own products, but that clearly isn't available to everyone — it eats your evenings and weekends, and for a long stretch nobody pays you a bonus for it. I don't have an answer to that one. If you do, I'd like to hear it. --- # The Greatest Product Managers, No. 10: Shigeru Miyamoto — He Handed Over Mario's Day-to-Day and Kept Exactly One Thing: He Plays the First 30 Minutes of Every Game Himself URL: https://doaipm.com/en/blog/miyamoto-first-thirty-minutes/ Published: 2026-08-12 Tags: Shigeru Miyamoto, Nintendo, Mario, Zelda, Product Manager Rankings, Product Decisions, Tech Commentary Shigeru Miyamoto is 72. His title at Nintendo is Executive Fellow and Representative Director. He has formally stepped back from day-to-day development on the Super Mario series and handed it to a younger generation of developers. But he kept one action: **he personally plays roughly the first 30 minutes of every game.** And after playing it, there is only one thing he is judging — whether it feels like Mario. ## Those 30 minutes are what the word "taste" actually looks like In this profession, talk about taste usually ends up empty. Some people call it aesthetics, some call it detail, some say it's the thing you can't articulate but know when you see. None of those can be handed over, and none of them can be checked. Miyamoto's 30 minutes are different, because they are executable: **a specific duration, a specific action, a specific test.** And the position is well chosen. The opening 30 minutes is the only stretch of a game that every single player experiences — however low the completion rate, first-half-hour reach is close to 100%. It is also the hardest part to get right: tutorial, pacing, the feel of the first jump, the timing of the first enemy, all crammed into the same window. Someone who hands off everything else and keeps only this segment knows exactly where his judgment is worth the most. ## Ranked 10th: taste 99 and originality 99 are top of the board, business 88 is the lowest in the top ten I had [Claude score "the 100 product managers who changed the world"](/en/rankings/). Miyamoto comes 10th, OVR 95, across six dimensions: | Dimension | Score | |---|---| | Vision | 95 | | Insight | 98 | | Taste | 99 | | Business | 88 | | Scale | 92 | | Originality | 99 | Only three people on the entire board score 99 on taste: Jobs, Allen Zhang, Miyamoto. Originality 99 is also top of the scale. On the other side, **business 88 and scale 92 are the two lowest numbers in the entire top ten.** Yesterday's piece was Gates — business 99, scale 99, with originality 94 the lowest in the top ten. Today's man is the exact inverse. They sit 8th and 10th on the same board, both at OVR 95, and the things that add up to that 95 are almost perfectly opposed. That is not a coincidence. It is the two poles of this profession: **one man put other people's inventions on every device; the other invented new things and never chased getting them onto every device.** ## Business 88 isn't a knock — it's the fights Nintendo chose not to have Nintendo has sat out round after round of competition. It didn't chase raw performance. It didn't chase the graphics arms race. It didn't chase open-world technical benchmarks. It moved into mobile slowly and warily, and third-party ecosystems have never been its strength. Those are real commercial costs. Sony and Microsoft left it several generations behind on hardware specs, and its console volumes, revenue base and platform take-rate are not in the same weight class. Miyamoto has said in interviews that this is no longer an era where customers chase specs alone. You can hear that as self-consolation, or you can hear it as a choice that has held for forty years. What he and Nintendo keep picking is the same thing — **between "how powerful is this machine" and "is this half hour fun," always bet on the second.** Business 88 is the invoice for that choice. Originality 99 and taste 99 are what it bought. ## What he's doing in 2026: taking the characters off the console Since stepping back from daily development, his centre of gravity has visibly moved. *The Super Mario Galaxy Movie* arrives in April 2026 with him as producer, in the final stage of production with Illumination. The live-action *Legend of Zelda* film, directed by Wes Ball, also has him as producer. In July this year he said publicly that he wants to keep taking Nintendo's IP beyond games, and that he wants to "bring new characters to the world." In a *Famitsu* interview he talked about Nintendo planning for a future where one household owns multiple consoles, so developers can build for Switch and Switch 2 at the same time. A man who at 72 handed over the day-to-day of the series he created, then poured his energy into cinemas, theme parks and new characters — **he is still doing the same job: putting a feeling in front of more people. Only the carrier changed, from cartridge to screen.** Worth noting: in the top ten of this board, Gates is the only one still writing on his own personal website, Page and Brin's homepages froze in 1998, and Miyamoto has no personal site at all. He doesn't write. He makes things. ## The thing you keep when you hand over is the judgment you believe can't be delegated Read the last two days together and a useful comparison appears. Gates handed over Microsoft's day-to-day and kept **fixing a 2045 closing date for the foundation** — what he considers undelegatable is endgame design and distribution. Miyamoto handed over Mario's day-to-day and kept **the first 30 minutes of every game** — what he considers undelegatable is whether the thing tastes right. Both let go. And the one item each of them held back says more accurately who they are than any single decision across forty years. For an ordinary product manager that turns into a directly usable question: **if you had to hand over what you're holding tomorrow, which one action would you keep for yourself?** That action is your actual job. Everything else can be hired. ## I can't name an action like that I thought about my own answer and found I don't have one. If I handed over what I build tomorrow, I could not name a kept action as specific as "the first 30 minutes." Everything I can think of comes out as "steer the direction" or "hold the quality bar" — phrases that cannot be handed over. And the reason they're empty is that they can neither be executed nor checked. Miyamoto's 30 minutes hold up because they can be put on a calendar, because other people can see him doing it, and because at the end he can give a yes or a no. A 72-year-old, stepped back from the day-to-day of the series he created, keeps playing the first half hour of every game himself. One test only: does it feel like Mario. (All scores and rankings in this piece were produced by Claude, an AI; methodology is on the rankings page.) --- # The Greatest Product Managers, No. 8: Bill Gates — He Wrote an End Date for His Own Organization: Close at the End of 2045, After Spending $200 Billion URL: https://doaipm.com/en/blog/gates-wrote-an-end-date/ Published: 2026-08-11 Tags: Bill Gates, Microsoft, Gates Foundation, Product Manager Rankings, Product Decisions, AI, Tech Commentary On January 14, 2026, the Gates Foundation's board approved a budget: $9 billion a year. The same announcement carried something less widely repeated — that number is the end point of a four-year plan, and the plan itself points at another date: **the end of 2045, when the foundation closes permanently.** Before it closes it will pay out another $200 billion, double what it spent in its first 25 years. Over the next five years it will cut roughly 500 roles. Not for lack of money. Because it has to finish on schedule. ## How rare it is to write an end date for your own organization Almost every organization's implicit goal is to survive. Budgets run annually, targets climb annually, teams grow with demand, and nobody writes "this will terminate in year N" into a proposal. Charitable foundations are the extreme case: most are designed to exist forever, living off investment returns while the principal is never touched. The Gates Foundation runs the opposite way: **spend the principal, then close, with the date fixed.** For a product manager that is a concrete technical move, not a posture. Once the termination date exists, the arithmetic of everything changes: - The budget stops being "what can we spend this year" and becomes years-remaining-until-2045 divided into what must be spent annually - Hiring stops being "how big should the team be" and becomes "what happens to these roles in the closing year" - Projects get filtered not by "is this sustainable" but by "does it finish inside the window" Those 500 roles are one output of that arithmetic. An organization with no end date does not lay people off while sitting on a $9 billion annual budget. ## Ranked 8th: business 99 is the board's top bracket, originality 94 is the lowest in the top ten I had [Claude score "the 100 product managers who changed the world"](/en/rankings/). Gates comes 8th, OVR 95, across six dimensions: | Dimension | Score | |---|---| | Vision | 97 | | Insight | 90 | | Taste | 84 | | Business | 99 | | Scale | 99 | | Originality | 94 | Business 99 and scale 99 are top of the entire board. Only two people on the whole list score 99 on business: him and Bezos. Originality 94 is **the lowest in the top ten**, and taste 84 sits at the bottom too, tied with Musk. The distribution is blunt: he was not the inventor. MS-DOS was bought. The graphical interface wasn't his. The browser was a scramble forced by Netscape. He was late to search and late to phones. The people scoring 99 on originality are Jobs, Musk, Miyamoto, Ford — the ones who invented categories. Gates being five points below them there is accurate. He won in the other two columns. ## Fifty years, one move: distribution Microsoft's early mission was a computer on every desk. Plenty of people said it; he did it — **and the way he did it was not by inventing the computer, but by making sure every computer other people built had his software inside.** The product was distribution itself. Licensing terms, OEM preinstalls, the developer ecosystem, file-format compatibility — none of that is invention. It is the engineering of getting one thing into every entrance. The five-point gap between originality 94 and business 99 is exactly the shape of this person. What he is doing in 2026 is the same move. The Gates Foundation and OpenAI built a program called Horizon1000, starting in Rwanda, applying AI to African health systems. The foundation has committed $1.4 billion to give farmers on the front line of extreme weather better information on weather, prices, crop disease and soil. In education, AI-driven personalized learning is now the largest line in the foundation's education spending. He didn't invent the diagnostic models, and he didn't train the language models. What he is doing is getting them into clinics in Rwanda and onto farmers' phones — **the same act as getting Windows onto every clone thirty years ago.** ## gatesnotes.com: the only one in the top ten still writing on his own site Writing about Page and Brin yesterday, I quoted their Stanford homepages, frozen since 1998, still listing them as Ph.D. Students. Gates is the only person in the top ten of this board still publishing on his own personal website. This year's outlook on `gatesnotes.com` is titled *The Year Ahead 2026*, and contains this line: > Of all the things humans have ever created, AI will change society the most. In the same piece the optimism comes with conditions. He calls 2026 a year that should be used to prepare for changes in the labour market, and flags the risk of AI landing in the hands of bad actors. A 70-year-old whose organization already has a closing date is still writing, one post at a time, on his own domain. That is consistent with his scores: he doesn't invent new forms, he keeps doing one thing for a very long time. ## My own end date is blank Nothing I build has a termination date written on it. Thinking about it, that isn't because I plan to work on these forever. It's because I have never been forced to answer the question — if this must end in a specific month of a specific year, what do I do this year, and what do I not do. Without that date, every trade-off can be deferred, and deferring requires no judgment at all. The hardest part of that budget isn't the $200 billion. It's that **he fixed the date first, and then let every number grow backwards out of it.** The $9 billion annual payout and the 500 role cuts are consequences of the date, not independent decisions. The budget approved on January 14, 2026 wrote this organization's remaining life as 19 years. Every dollar it spends now is counted backwards from that end. (All scores and rankings in this piece were produced by Claude, an AI; methodology is on the rankings page.) --- # They Organized the World's Information. Their Own Homepages Stopped in 1998 — and Still Say They're 'Currently Working on Google, a Search Engine for the Web' URL: https://doaipm.com/en/blog/the-homepage-they-never-updated/ Published: 2026-08-10 Tags: Larry Page, Sergey Brin, Google, Gemini, Product Management, Product Decisions, Tech Commentary On the morning of August 10, 2026, I pasted these two addresses into a terminal: ``` http://infolab.stanford.edu/~page/ 200 http://infolab.stanford.edu/~sergey/ 200 ``` Both alive. Two graduate students' homepages, still sitting on a Stanford server, twenty-eight years on. ## Page's says "Ph.D. Student" and gives a 1998 office phone number The page is titled "Lawrence or Larry Page's Page." What follows is copied verbatim: > Larry (Lawrence) Page > Ph.D. Student > Computer Science Department > Stanford University > > Member of Terry Winograd's Project on People, Computers, and Design. > > **Currently working on Google, a search engine for the Web.** Papers and a demo are available off this page. > > email: page@cs.stanford.edu > office: Gates 360 > phone: (650) 330-0100 Present tense. Papers and a demo are available off this page. The office number and the landline are still there. Brin's page reads the same way: > Sergey Brin's Home Page > Ph.D. student in Computer Science at Stanford — sergey@cs.stanford.edu > > Research > **Currently I am at Google.** > In fall '98 I taught CS 349. Below that is his data mining work and a list of papers. One of them, co-authored with Lawrence Page, is titled *Dynamic Data Mining: A New Architecture for Data with High Dimensionality*, and carries this note: > We describe a new architecture for data mining (sorry not yet available online). **Work in progress.** Another is marked "To appear in VLDB '98." ## That sentence was worth one line, next to a data mining paper The tone of those two pages is the whole point. "Currently working on Google, a search engine for the Web" carries exactly the same weight on that page as "In fall '98 I taught CS 349." A doctoral student's status list: I'm with Terry Winograd's group, I taught a course in the fall, I do data mining, I'm building a search engine. No vision. No mission. No organizing the world's information. That line came later. **The thing product managers most reliably overestimate is their own description of a project on the day it starts.** We're trained to write down what this will change, as if failing to say it means we haven't thought hard enough. And the thing that actually changed the world was, on the day it began, worth one line on its author's own homepage — sitting beside a fall course listing. This is not an argument that vision is useless. It's that **how you narrate something at the start has almost no relationship to how large it turns out to be.** Twenty-eight years later, the restraint of that one line reads as accurate: in 1998 Google really was a demo and a few papers. ## Ranked 7th: business 98, scale 99, insight only 90 I had [Claude score "the 100 product managers who changed the world"](/en/rankings/). Page and Brin share one entry at 7th, OVR 95, across six dimensions: | Dimension | Score | |---|---| | Vision | 96 | | Insight | 90 | | Taste | 86 | | Business | 98 | | Scale | 99 | | Originality | 97 | Scale 99 and business 98 sit in the top bracket of the whole board. One input box absorbed humanity's demand for information and defined the internet's business model for twenty years along the way. Nothing to argue with there. The other two are conspicuously low. **Insight 90 and taste 86 both rank last within the top ten.** Allen Zhang's insight is 99. Miyamoto's taste is 99. Jobs is above 98 on both. That distribution says something quite specific: they did not win by understanding users better than anyone else. PageRank grew out of a paper — ranking pages by the citation structure of links is an academic judgment. The product decision they actually got right was **turning an academic algorithm into a business**, and getting it onto every device on earth. In this profession those are two different talents. ## Twenty-eight years on: Brin is in the office almost daily, Page went to the factory The two people who wrote those pages have gone in different directions. Brin stepped back from Alphabet's day-to-day operations in December 2019, then came back. He has said publicly that staying retired would have been a "big mistake," that during that stretch he felt he was "spiralling" and "a bit less sharp." By 2023 he was at the office three or four times a week, working alongside researchers. More recently he says he is at Google "pretty much every day now," helping train the latest Gemini models. He has also said, more bluntly, that it is time for retired computer scientists to get back to work. A 52-year-old who never needs to earn another dollar went back to do something as concrete as training models. Page did not go back. He is building a company called Dynatomics — AI generates optimized product designs, factories build them. It is led by Chris Anderson, formerly CTO of Kitty Hawk, the flying car company Page also backed. On December 30, 2025, he filed to move the company's registration from California to Keller, Texas, while leasing 73,400 square feet of R&D space in Palo Alto, keeping the core engineering there. On April 30, 2026, after Alphabet reported surging cloud revenue, Forbes' real-time list pushed Page's net worth past $300 billion, second in the world. One went back to the place his 1998 homepage still describes. The other went off to do something else entirely. And both pages still say Ph.D. Student. ## All I did today was run two curls What I build is several orders of magnitude smaller, but the feeling transfers. Open the earliest README of your own project and you'll find phrasing that was meant to be temporary — "for now," "we'll handle that case later." Some of it got fixed. Some of it stayed, because it turned out to be accurate the whole time. Page's page may still be up simply because nobody maintains that server. But the effect of reading it is something else entirely: **it preserves, intact, what a thing looked like at the beginning — including how much its author underrated it.** We write project proposals big because we're trying to convince someone. Those two pages are the proof that how you describe it on day one and what it becomes are two independent lines. On August 10, 2026, both addresses still return 200. They say Ph.D. Student, Gates 360, (650) 330-0100, and one sentence in the present tense about currently working on a search engine for the Web. (All scores and rankings in this piece were produced by Claude, an AI; methodology is on the rankings page.) --- # How Expensive Can One Product Judgment Get? Zuckerberg Renamed the Company Meta for the Metaverse, Burned $80 Billion in Four Years, and Is Now Cutting the Budget URL: https://doaipm.com/en/blog/zuckerberg-the-name-you-cant-roll-back/ Published: 2026-08-09 Tags: Mark Zuckerberg, Meta, Metaverse, Reality Labs, AI Glasses, Product Management, Product Decisions, Tech Commentary In 2025, Meta's Reality Labs booked $2.21 billion in revenue and an operating loss of $19.19 billion. It lost $8.70 for every dollar it took in. The fourth quarter alone lost $6.02 billion on $955 million in sales — the worst quarter the division has ever had. ## What he said the day of the rename On October 28, 2021, at the Connect conference, Zuckerberg announced that Facebook's parent company would be renamed Meta. His words: "From now on we're going to be the metaverse first, not Facebook first." Meta comes from the Greek for "beyond." The same event announced something else: 10,000 hires in Europe over five years to build the metaverse. Renaming the company is the most expensive commitment a product manager can make. A release can be rolled back, a feature can be pulled, a company name cannot — it is simultaneously written into the legal entity, the domain, the sign on headquarters, every contract, every news story, every employee's business card. Carve a judgment into that layer and you have announced that there is no way back. ## The MVRS ticker they promised never traded a single day That rename announcement contained a line most people skip: the stock ticker would change from FB to **MVRS** — short for metaverse — effective December 1, 2021. That symbol never appeared in a single day's quotes. On May 31, 2022, Meta put out a new release: from June 9, its Class A common stock would trade on Nasdaq under **META**, replacing the FB ticker it had used since the 2012 IPO. So what actually happened is this: the company name became Meta, and the ticker did not become MVRS. The former points at "beyond" — abstract enough to be reinterpreted later. The latter would have nailed the four letters of *metaverse* into the trading symbol itself, and there is no reinterpreting that. That was the first place in this whole story where an escape hatch was left open. Almost nobody noticed at the time. ## Five years of losses: from $10.2B to $19.2B, never once narrowing Reality Labs operating losses by year: | Year | Operating loss | |---|---| | 2021 | $10.2 billion | | 2022 | $13.7 billion | | 2023 | $16.1 billion | | 2024 | $17.73 billion | | 2025 | $19.19 billion | Every year after the rename lost more than the year before. **Not one year narrowed.** Counting from the end of 2020, the division has accumulated close to $80 billion in operating losses. On the Q4 2025 earnings call, Meta's CFO said Reality Labs operating losses in 2026 were expected to stay near 2025 levels. The shape of that curve is itself a verdict. This is not "the investment phase hasn't reached the harvest phase" — an investment phase is supposed to widen and then narrow. Five straight years of widening only says something else. ## Ranked 9th, vision 94: the lowest in the top ten I had [Claude score "the 100 product managers who changed the world"](/en/rankings/). Zuckerberg comes in 9th with an OVR of 95, across six dimensions: | Dimension | Score | |---|---| | Vision | 94 | | Insight | 96 | | Taste | 84 | | Business | 96 | | Scale | 99 | | Originality | 95 | Scale 99 is not controversial — three billion people connected into one social graph, a magnitude only a handful on the entire board have reached. Insight 96 holds up too: the News Feed changed how humanity consumes information, and the Instagram and WhatsApp acquisitions were both called wildly overpriced at the time. Nobody says that now. The number worth talking about is vision 94. It is the lowest in the top ten — five points below Musk's and Jobs's 99, and one point below Miyamoto, who ranks behind him. Vision is defined as seeing a future others cannot. Zuckerberg has done it: the pivot to mobile was fast, brutal and textbook. With the metaverse the problem is not that what he saw was fake — AI glasses selling well today is precisely the evidence that "computing eventually leaves the phone and moves onto your face" was the right direction. The problem is timing. He took a judgment that may need fifteen years and bet it at three years' worth of confidence — and bet it on the company name. **Vision is never marked down for seeing the wrong direction. It is marked down for getting the timing wrong and betting as if you hadn't.** ## Cut 30%, while every other department was asked for 10% In December 2025, Bloomberg reported, citing people familiar with the discussions, that Meta executives were weighing cuts of up to 30% to the metaverse division in the 2026 budget. The sharpest detail in that report is the contrast: Zuckerberg asked **every** department to find 10% in cost savings, and told the metaverse team to go deeper. The cuts would include layoffs, falling hardest on the VR group, with Horizon Worlds also on the list. Execution followed. In January 2026, Reality Labs cut more than a thousand roles; in March it cut hundreds more. The savings have a clear destination. Meta's 2026 capital expenditure guidance is $115 billion to $135 billion, close to double the prior year, and the hardware line shifted from VR headsets to AI glasses. ## The new bet is phrased exactly like the old one On the Q4 2025 earnings call, Zuckerberg said "we're at a moment similar to when smartphones arrived," and "It's hard to imagine a world in several years where most glasses that people wear aren't AI glasses." Put that next to the 2021 line: > 2021: From now on we're going to be the metaverse first, not Facebook first. > 2025: It's hard to imagine a world in several years where most glasses that people wear aren't AI glasses. Same man, same sentence shape, different noun. The difference is that this time he did not rename the company. ## The cost of retreat depends on which layer you bound it to The budget can be cut 30%. A thousand roles can go. Three VR studios can close. The hardware line can move from headsets to glasses. All of it is reversible, because all of it is bound to the resource-allocation layer — and resource allocation is by nature redone every year. The company name is not reversible. Meta's AI glasses, its superintelligence lab, its $135 billion of capex all hang today under a sign that says *metaverse*. That division's own budget is being cut by a third, and the company still carries the name. When product managers make a judgment, they usually price one thing: how long it would take to undo if it's wrong. The thing actually worth pricing is the second: **which layer am I writing this into.** Into one release, and undoing it is a rollback. Into a public commitment, and undoing it is an apology. Into a product name and a URL, and undoing it is months of redirects. Into the company name, and it doesn't undo. The same wrong judgment can differ in cost by three orders of magnitude, and the difference isn't in the judgment. It's in how deep you drove it. ## The most expensive bet I ever made was writing a judgment into a URL What I build is much smaller than this. The most expensive bet I made was putting a feature's name into the title of a public doc and into the URL path, because at the time I was sure the concept would hold and deserved a formal name. Six months later the concept turned out to be wrong — not that the feature worked badly, but that the name I gave it drew the wrong boundary. Users read the name, understood it accordingly, and used the thing sideways. Changing the code took a day. Changing the URL and every reference to it in the docs took two weeks. To this day a few external links still point at the old path and land on a 404. I can't edit someone else's blog. Eighty billion dollars and two weeks are orders of magnitude apart, and they are the same kind of thing: whichever layer you carve a judgment into is the layer where you pay to remove it. And the mistake people make most often is carving it one layer deeper than necessary at the exact moment they feel most certain. On the ticker, back then, Zuckerberg held something back. In 2025 that division took in $2.21 billion, lost $19.19 billion, and is having up to a third of its 2026 budget cut. The company is still called Meta. (All scores and rankings in this piece were produced by Claude, an AI; methodology is on the rankings page.) --- # Your Hardware Just Shipped a Serious Defect — Now What? The GAC Toyota bZ7 Was Recalled Twice in Under 100 Days, Both Times for Software URL: https://doaipm.com/en/blog/ota-recall-is-still-a-recall/ Published: 2026-08-08 Tags: Vehicle Recalls, OTA, Software-Defined Vehicles, GAC Toyota, Product Management, Quality, Tech Commentary On August 1, GAC Toyota filed two recall notices. The first, S2026M0083V, covers 15,266 bZ7s. The stated cause reads: the smart Bluetooth module's software control program was insufficiently considered, and under certain conditions may send an abnormal command to the vehicle control module, causing the car to shift automatically from D to N while driving. The second, S2026M0084V, covers 24,286 bZ7s. The stated cause: the thermal management controller's software control strategy is incomplete, and after the vehicle powers on, the A/C compressor and water heater may stop working. In extreme cases defrost and defog performance degrades, obstructing the driver's view. That is 39,552 cars. The bZ7 had been on sale for less than 100 days. The remedy on both notices is the same sentence: a free over-the-air software update, no dealer visit required. ## The other three notices that month need hands on metal From August 8, two Jaguar Land Rover recalls take effect together. S2026M0080V covers 76,422 imported Range Rover and Discovery vehicles; S2026M0081V covers 83,810 Defenders. Total: 160,232. The defect is fretting corrosion at the connector between the steering wheel clock spring and the driver's airbag, raising circuit resistance and potentially preventing the driver's airbag from deploying. The remedy is applying a special grease to the connector terminals, free of charge. From August 28, S2026M0082V recalls 33,473 XPeng X9s. Manufacturing process variation reduced the airtightness of the front air springs; after long use in hot, humid conditions they may leak slowly, and in extreme cases handling is affected. The remedy is a free replacement of the improved front air spring strut assembly. Put all five side by side: | Recall number | Brand / model | Units | Where the defect lives | How it gets fixed | |---|---|---|---|---| | S2026M0080V | Range Rover, Discovery | 76,422 | Airbag connector fretting corrosion | Grease each car by hand | | S2026M0081V | Land Rover Defender | 83,810 | Same | Grease each car by hand | | S2026M0082V | XPeng X9 | 33,473 | Front air spring airtightness | Replace strut assembly per car | | S2026M0083V | GAC Toyota bZ7 | 15,266 | Bluetooth module software | Push one OTA | | S2026M0084V | GAC Toyota bZ7 | 24,286 | Thermal controller software | Push one OTA | The top three rows mean more than 190,000 cars driving into service bays one at a time — lifts, labor hours, parts, customer appointments, all of it real. The bottom two rows are 39,552 cars and an engineer publishing an update package. ## The wording of defect causes is changing: from "corrosion" to "insufficiently considered" Read through two months of recall notices and the vocabulary of classic hardware defects is a fixed set: corrosion, metal fatigue, cracking, insufficient design strength, half-shafts of the wrong size fitted to the front axle, lithium growth inside the cell. Harley-Davidson's batch reads "the upper triple clamp material strength is insufficiently designed, and cracks may form at stress concentrations." Volvo's EX30 batch reads "production variation may cause lithium growth inside the cell." The new vocabulary is a different set: insufficiently considered, incomplete strategy, improperly configured. In the same run of notices, FAW Toyota recalled 9,408 Prados because "the combination meter control program was improperly configured, and at vehicle start the combination meter may fail to start correctly." "The smart Bluetooth module's software control program was insufficiently considered," translated into what a software team says every day, means a boundary condition never made it into the check. In a requirements review that is one line of minutes reading "let's not handle that case for now." In a recall notice it is 15,266 cars. ## Of 1,904 OTA updates in 2025, only 13 counted as recalls SAMR's 2025 figures: 190 vehicle recalls nationwide covering 6.846 million vehicles, down 18.5% and 39.1% year over year. Of those, 13 recalls were carried out over the air, covering 1.756 million vehicles — 25.7% of all vehicles recalled that year. In the same year, manufacturers reported 1,904 OTA updates covering 140 million vehicle-instances, with the number of updates up 38% from the year before. 1,904 and 13. The overwhelming majority of OTAs add features, tune the experience, fix small annoyances. Thirteen were classified as recalls. Where that line gets drawn is the most expensive judgment in this whole business — on one side is product iteration, on the other is a defect recall that must be filed with the regulator, that stops production and sales, and that lands in the annual statistics. The March 2025 joint notice from MIIT and SAMR draws it explicitly: "Where an enterprise carries out an OTA update to eliminate a defect in an automotive product and implement a recall, it shall organize the recall in accordance with the Implementing Measures for the Regulations on the Recall of Defective Automobile Products, and immediately halt production and sale of the defective product." The same document adds a line aimed squarely at the temptation: standardize how OTA updates are applied, so that enterprises do not use them to conceal vehicle defects or evade responsibility. > OTA has crushed the cost of a fix from "190,000 cars through service bays" down to "publish one package," and it has not removed a single filing obligation. Being able to change it remotely does not mean you may change it quietly. ## After a defect is confirmed: ten working days, three months, 1% to 10% of value A recall does not start on the day the notice goes public. It starts on the day someone concludes a defect may exist. For vehicles, the Regulations on the Recall of Defective Automobile Products require the manufacturer to file a recall plan with the competent authority before carrying out the recall. Hardware that isn't a car falls under the Interim Provisions on the Administration of Consumer Product Recalls, which spell the clock out in more detail. Article 9: where a producer discovers that a consumer product may be defective, it shall immediately organize an investigation and analysis. Article 17: for a voluntary recall, the producer shall report the recall plan to the provincial market regulator within ten working days of the investigation concluding that a defect exists. Article 21: an interim summary every three months from the start of the recall, and a final summary within fifteen working days of completing the plan. The penalty column differs by an order of magnitude. Article 24 of the vehicle regulation sets fines at 1% to 10% of the value of the defective automotive products, with license revocation in serious cases; Article 23 sets 500,000 to 1,000,000 yuan for failing to cooperate with a defect investigation and refusing to correct. Article 25 of the consumer product provisions sets 10,000 to 30,000 yuan for failing to correct in time. Which makes the hardest move in the chain not the fix but the **classification** — whether anyone will call it a defect. Of the 190 recalls in 2025, 55 were carried out under SAMR prompting, covering 2.934 million vehicles, or 42.9% of all vehicles recalled that year. By unit count, nearly half the cars were not raised by the manufacturer's own hand. The fix, the public statement and the compensation package can all be ready inside 48 hours. What stalls is always the step before: someone has to sign off that this is a defect. ## On my side, the worst case is a rollback I build software, not hardware. When something breaks in production, the worst night I have had ends with rolling back a release and writing the postmortem the next morning. The cost to users is a few hours of not being able to use the thing, or a batch of data that has to be reimported. I can afford that, which is why "ship it and fix it later" almost never required hesitation from me. The bZ7 notice says the car may shift automatically from D to N while driving. Same species of mistake — a control path with a case left out — and on my side it ends in a rollback, on theirs someone loses power on the road. Once software defines hardware, the worst outcome of one commit goes from "the user restarts the app" to a public notice with a filing number, a unit count and a regulator's stamp. And the people writing that logic still carry software habits: ship, watch the data, iterate. OTA makes that habit harder to break. The more a fix looks like a release, the easier it is to judge "is this even a problem" by release standards. The bZ7 filed two recalls in under 100 days, and the earliest production date in scope is October 28, 2025 — both defects left the factory with the first cars, and were simply driven around for a few months before anyone owned up to them. The five notices of August 2026, numbered S2026M0080V through S2026M0084V, cover 233,257 vehicles. For 39,552 of them, the problem is a condition someone left out of the code. --- # Is a PM's Record Measured by the Exit or by Survival? Musk's X Product Head Sold Two Hits — Both Were Shut Down URL: https://doaipm.com/en/blog/sold-or-alive/ Published: 2026-08-07 Tags: Nikita Bier, X, Elon Musk, Product Management, Careers, Tech Commentary "Time to pass the torch and demote myself to my natural state: a poster." That's what Nikita Bier wrote on August 5, stepping down as head of product at X. He took the job in July 2025, lasted a year and one month, and stays on as an adviser. A man who ran product for a platform with hundreds of millions of users says his natural state is posting. You can read that as self-deprecation. You can also read it as unusually honest self-knowledge. ## His record is strong, but strong in one particular way Bier built two products before this. One was tbh, an anonymous polling app that caught fire among American teenagers and sold to Facebook in 2017. The other was Gas, an anonymous compliments app — same audience, same explosion — sold to Discord in 2023. Two startups, two exits. That record stands up anywhere in Silicon Valley; most people never manage it once. TechCrunch added one line when reporting his departure: both apps were later shut down. tbh went into Facebook and was killed. Gas went into Discord and was killed. Nothing he built is still running today. ## Selling and surviving are two different objective functions This isn't nitpicking. It's a question of what his operating system was calibrated for. Teen social products have a well-worn playbook: find a mechanic that spreads inside a school, ignite it as fast as possible, and sell while the heat is still on. Every step of that playbook optimises for the exit — the growth curve has to be steep, the users have to be concentrated in a demographic an acquirer can recognise, and the product doesn't need a story for its second year. It actively discourages you from thinking about year two. Thinking about year two means retention, monetisation and content moderation, and all of those flatten the curve you're about to show a buyer. By that objective function, Bier scored full marks twice. The product being shut down was never his KPI — his KPI settled the moment the deal closed. The trouble is that when this system leaves the startup environment, the scoring rubric changes. ## Thirty new products in a year is itself the answer Bier's own summary of his year at X: 30 new products shipped, while "protecting the integrity of the town square." X's account adds that he rebuilt the feed, the Android client, DMs and notifications, and delivered X Money — the payments feature Musk had promised in 2024 and never shipped. X Money is hard work, and landing it is real. So are the feed and DM rebuilds. But "30 new products in a year" is a strange number for a platform that already has hundreds of millions of users. A startup trying 30 things in a year is normal — you're hunting for the mechanic that spreads, and the cost of a failed experiment is roughly zero. A mature platform shipping 30 new products in a year means most of them were never really raised. The hard part on a platform product has never been inventing features; it's getting one feature to survive its first quarter, enter a habit, and not get crowded out by next quarter's batch. **Casting a wide net for a hit and raising one thing into a habit are two different crafts.** He is plainly world-class at the first. ## So how should a product manager's résumé be scored That's what I actually found myself thinking about. The industry's default is to score by exit: how many companies you sold, for how much, to whom. There's logic to it — an exit is verifiable. It has a number, an announcement, a date. Unlike "retention," it can't be narrated into anything you want. It has one side effect: **it removes "did the product survive" from the evaluation entirely.** Scored by exits, Bier is two for two. Scored by survival, he's zero for two. Neither scoring is complete. Scoring only by survival would be unfair to a lot of people — products get shut down for reasons that have nothing to do with their founder. Big companies buy you for the people and the technology, not to keep the app running forever. tbh and Gas being killed is mostly on Facebook and Discord, not on Bier. But scoring only by exit produces an awkward kind of career: a person can hold a genuinely impressive résumé while nothing he ever built still exists. ## He answered it himself Bier didn't say he was off to build the next hit, and didn't say he was returning to Lightspeed to invest. He said he was going back to posting. **That line is more honest than any retrospective.** He knows what he is best at and enjoys most — spotting a gap in attention right now and turning it into a moment. Posting and building a viral app use the same muscle. Running a platform that hundreds of millions of people open every day uses a different one. ## What's left X has split the product lead across three people: Benji Taylor on design, Jonah Katz on iOS, Mridul Singhai over the product organisation. Splitting one person's job three ways says something on its own — either that role should never have sat on one person, or X wants to play it differently. As for how a product manager's record should be scored, I don't have an answer. I only know that when I next read the line "serial entrepreneur, two successful exits," I'll now ask one more question: are those two products still running? --- # AlphaGo's Demis Hassabis Steps Back From Running DeepMind Day-to-Day — and Google's Chief Scientist Walks Out the Same Day With Three Others URL: https://doaipm.com/en/blog/jeff-dean-operating-system/ Published: 2026-08-06 Tags: Jeff Dean, Google, DeepMind, Demis Hassabis, AI-Era Product Management, Tech Commentary On August 5, Google announced that Demis Hassabis is becoming chair of Google DeepMind and handing off day-to-day operations. Alphabet fell more than 5% that day. Hassabis needs no introduction. AlphaGo beating the world's best Go players was a global moment; AlphaFold later won him a Nobel Prize in Chemistry. So when the news broke, everyone talked about him. The same day carried another item that almost nobody looked at closely: Jeff Dean, Google's chief scientist, is leaving to start a company. The 5% drop is a result. Hassabis moving up is only half a signal. The half that matters is inside the item nobody read. ## Two men who shared one computer left together Jeff Dean joined Google in 1999 as employee number 30 and stayed 27 years. If the name doesn't register, that's normal — the things he built are things you use every day, but none of them has his name on the outside. Leaving with him is Sanjay Ghemawat. That name has almost no public profile, but inside Google he and Jeff Dean are the only two level 11 Senior Fellows the company has ever had — the highest technical rank there is, awarded to exactly two people in twenty-odd years. In December 2018 *The New Yorker* ran a piece about them called "The Friendship That Made Google Huge." It isn't really about their achievements; it's about the fact of two people writing code hunched over the same computer. Dean says in it that what you want is a pair-programming partner whose way of thinking is compatible with yours, so that the two of you together are a complementary force. Also leaving: Quoc Le, a founding member of Google Brain, and Oriol Vinyals, a senior research scientist at Google DeepMind. Four people, one new company, called Discovery Loop, with Jeff Dean as CEO. ## He spent 27 years doing the same thing Lay his résumé out in a line and a very monotonous pattern shows up. Early on he scaled Google's search index by 100× and its query-serving capacity by 1000×; Google's first ads system and content ads platform were his. Then came MapReduce, BigTable, Spanner, TensorFlow, Pathways — a string of names that don't sound like products, because they aren't. They're the floor other people stand on to build products. The TPU program was his too. Later he led Google Brain and Google Research, working on word2vec and neural architecture search. Twenty-seven years, one repeated act: taking something so hard nobody could do it and turning it into something everybody can just use. That operating system carries an unstated premise: that you have time to build a foundation. You build it, others use it, you collect rent from the whole world. Nearly ten years separate MapReduce from TensorFlow. The TPU took years from kickoff to being genuinely good. This worked for Google's first twenty-five years because for those twenty-five years Google set the rules and held the clock. ## The same day, Google moved the clock forward Hassabis is becoming chair of Google DeepMind and also Alphabet's chief scientist. Jeff Dean's title on his way out was Google chief scientist — so on the same day that title changed hands and moved up a level, from Google to Alphabet. Hassabis put it this way: this is the right time to hand over the day-to-day running of Google DeepMind so that he has the time and space to focus on the bigger picture. Taking over is Koray Kavukcuoglu, previously chief AI architect, promoted to senior vice president and reporting directly to Sundar Pichai. One goes upstairs to think; one comes down to deliver. In fairness, this isn't a squeeze-out or a palace coup. Hassabis co-founded DeepMind and has run it since Google acquired the lab in 2014; if he were really being sidelined, it wouldn't look like this. What Google is admitting is that at this stage the scarce resource is not people who think about big questions — it's people who can ship every quarter. Model release cadence, enterprise contracts, distribution channels: that's what this year's AI race runs on. A senior vice president reporting straight to the CEO and accountable for delivery is more urgently needed than a chief scientist thinking about the bigger picture. The clock sped up. And nobody can wait for a foundation to be poured. ## Google didn't push him out — it found him somewhere that fits Discovery Loop is registered as a public benefit corporation and works on automating scientific research with AI. Its own framing: to use massive computational scale to fundamentally transform the speed and efficiency of innovation. Translated: build a foundation that makes scientific research faster. Which is what Jeff Dean has been doing for twenty-seven years. The only difference is that the foundations used to be for Google engineers, and now they're meant for scientists everywhere. Radical Ventures and Khosla Ventures are co-leading the round, with Kleiner Perkins, Lightspeed and Doerr Capital joining. There's one more name on the list: Alphabet put money in too. Pichai was explicit that Google will keep working with Discovery Loop as a founding investor and cloud partner. So the full shape of this is: Google sees its four strongest foundation-builders to the door, invests in them, and plans to sell them cloud. It isn't that Google couldn't keep him. It's that there was nowhere to put him. Google's organisation now turns on a quarterly cycle, and a foundation project that takes three to five years to pay off has no place inside it where it could even be chartered. Letting it out is cleaner — funded by outside capital, tethered by cloud, and when it finally builds something, Google is first in line to buy. ## That system now lives outside the building Google's first twenty-five years were built on Jeff Dean's operating system. The index, the ads, MapReduce, BigTable, the TPU — every layer stacked out of the slow, careful work he and Sanjay did. Now that system has been moved outside the company, and Google is still paying to keep it alive. Whether that's shrewd or nervous is too early to say. You'll know when Discovery Loop pours its first foundation — or when Google next needs something only Jeff Dean could build and finds he isn't in the building. --- # Why Is Google Guaranteeing $43.8B of Other People's Data Center Rent? TPU Operators Borrow at 7.1%, Nvidia's at 9.3% URL: https://doaipm.com/en/blog/anthropic-six-compute-suppliers/ Published: 2026-08-05 Tags: Google, Anthropic, Nvidia, TPU, Compute, Off-Balance-Sheet Financing, Product Management, Tech Commentary There's a number in Alphabet's Q2 2026 10-Q: as of June 30, the ceiling on third-party data center rent Google has promised to cover if tenants default stands at **$43.8 billion**. Nine months earlier it was $6.5 billion. A sevenfold increase. None of those data centers belong to Google. It guaranteed somebody else's rent. ## What $43.8 billion buys is 2.2 percentage points A guarantee costs nothing to issue. What it transfers is credit. With Alphabet standing behind them, lenders price these projects differently overnight: **operators running Google TPUs borrow at 7.1%, while comparable projects inside Nvidia's ecosystem pay 9.3%.** A structural 2.2 points. In an industry where capex runs to tens of billions and construction is heavily debt-financed, 2.2 points is not a finance detail — it decides whether a project gets built at all. To get a sense of scale: on a $10 billion ten-year facility, the interest gap is $220 million a year, north of $2 billion cumulatively — enough to eat a data center's operating profit across its entire life. So the nature of this should be stated plainly: **what Google is using against Nvidia is not chip performance. It's its own credit rating.** It didn't cut the price of TPUs. It cut the cost of borrowing for people who buy TPUs. Google has now extended this kind of guarantee across **10 TPU data center projects**, totalling **2.4 gigawatts** of power capacity. ## At the other end of the chain is a company that can't borrow The primary customer this structure serves is Anthropic. Its run-rate revenue just crossed **$30 billion** — at the end of 2025 that figure was around $9 billion. The growth is steep, but it is a private company with no credit rating, and infrastructure financing in the tens of billions is not something its own balance sheet can raise. Google's guarantee is what fills that gap. What flows along this chain is Alphabet's credit; what Anthropic supplies is demand and rent. The supply end is locked down too. In April 2026, Broadcom, Google and Anthropic expanded their arrangement to roughly **3.5 gigawatts** of compute, delivering from 2027; separately, Broadcom and Google hold a long-term TPU supply agreement running to **2031**. Analysts estimate Broadcom will book around $21 billion of Anthropic-related revenue in 2026 and about $42 billion in 2027 — those two figures are sell-side estimates, not disclosed in filings. ## Four parties in the middle, none of whom wants the hardware on their books - **Morgan Stanley** built a special purpose vehicle called Compute SPV, structuring the deal and packaging the guarantees into bond products for investors. - **Apollo and Blackstone** provide the money through private credit. The SPV uses that outside capital to buy chips and leases the hardware to Anthropic. **The first transaction, in June 2026, was $35 billion for roughly 1 million TPUs.** - **Broadcom** provides about **$30 billion** in residual value guarantees, the buffer for when Anthropic can't make lease payments. Its total procurement commitments to Google run to **$128 billion** ($55.2B in FY2027, $72.9B in FY2028). - **Crypto miners** — TeraWulf, Cipher, Hut 8 — supply grid power and sites. The **360 megawatt** facility TeraWulf built in upstate New York is dedicated to Anthropic; Google guarantees the lease and takes share purchase warrants in return. Back in October 2025, Morgan Stanley had already packaged TeraWulf's lease into construction bonds and raised **$3.2 billion**. There's one more layer: AI cloud provider Fluidstack leases and develops data centers packed with Google TPUs, then rents that compute on to Anthropic — with the leases, again, guaranteed by Google. Line the roles up and the common thread is obvious: **not one party wants the hardware recorded on its own books.** The chips are bought with outside investors' money, held by an SPV, hosted on miners' sites, backstopped by Broadcom's residual guarantee, and the rent is guaranteed by Google. Every link pushes the asset and the risk to the next one. ## What's actually scarce isn't chips — it's grid power The miners' presence in this network is the most revealing part of the whole thing. Chips can be ordered and capacity can be scheduled, but **hundreds of megawatts of grid interconnection cannot be conjured up**. What crypto miners have done for the past decade is find cheap power, secure the interconnection, and get machines energised as fast as possible. As mining economics deteriorated, the idle interconnection capacity in their hands turned out to be exactly what AI is short of. That also explains the deal on August 4 that looks most absurd on its face: Anthropic signed a **$10 billion, six-year** compute contract with Volta — a company **founded this year** — for a **133 megawatt** facility in Norway, with the actual construction handled by the crypto mining company **Bitdeer**. Norway has cheap hydro power and free cooling. Mining companies have an off-the-shelf ability to turn electricity into a machine room. This isn't an isolated oddity; it's the same logic replicated in another location. ## How much more is off the books The 10-Q discloses more than that $43.8 billion: | Item | Amount | |---|---:| | Recognised data center lease guarantees | **$43.8B** | | Data center leases not yet commenced | **$85.2B** | | Guarantees tied to power and energy facilities | $7.6B | | Planned guarantees, terms pending | $24.1B | Per reporting, the liability Google actually recognises on its balance sheet for these guarantees is only about **$815 million** — because accounting records the fair value of a guarantee, not the maximum exposure. None of this machinery is new to Wall Street; leasing and project finance have run on it for decades. What's new here is that **it's being used to sell a chip.** ## One buyer, three completely different ways of paying Anthropic gets compute from each supplier by a different mechanism, and the mechanism reveals what each seller is after. | | Google | Nvidia / Microsoft | SpaceX | |---|---|---|---| | Instrument | **Credit guarantee** | **Equity investment** | **Cash lease** | | Specifics | Rent guaranteed across 10 projects, 2.4 GW | Nvidia up to $10B, Microsoft up to $5B | Whole of Colossus 1 | | What the buyer gives | Rent | Committed $30B of Azure + up to 1 GW of Nvidia systems | $1.25B/month through May 2029 | | What the seller is after | Making TPU the alternative to Nvidia | Locking in demand and technical direction | Monetising idle capacity | | Where the risk sits | Off-balance-sheet, $43.8B exposure | On the balance sheet, as an investment | Nowhere — it collects cash | The Nvidia one is the most circular: on November 18, 2025, Nvidia announced an investment of up to **$10 billion** in Anthropic and Microsoft up to $5 billion, and on the same day Anthropic committed to purchase **$30 billion** of Azure compute plus up to 1 gigawatt of Nvidia systems. The money goes around and lands back on the seller's books. After that round Anthropic's valuation reached roughly **$350 billion**, against $183 billion in September 2025. The three instruments share one purpose: **keep this buyer able to buy, and buying.** They differ only in where each party is willing to park the risk. SpaceX also sells compute to Anthropic, and its deal is a completely different animal from Google's. After xAI and SpaceX merged, xAI moved its training to Colossus 2, leaving Colossus 1 in Memphis empty — over 220,000 Nvidia GPUs, over 300 megawatts, at roughly **11% utilisation**. It was then leased in its entirety to Anthropic for **$1.25 billion a month** through May 2029. Musk's response to the arrangement: "No one set off my evil detector." That is genuinely disposing of idle capacity: the facility is already built, it depreciates every day, renting it to anyone beats leaving it dark, and no financial structure is required. **Google's version isn't disposing of idle capacity — it's manufacturing capacity.** One monetises a sunk cost; the other mobilises outside capital to build new supply and then distributes the new risk. Both look like "selling to a rival." Underneath they are not the same thing at all. ## The path the risk travels back The whole structure rests on one condition: that Anthropic can pay the rent. It's at a $30 billion run rate now and growing fast. But if demand falls short, or competition erodes its pricing power, pressure moves back up the chain in order: 1. Anthropic can't make the payments 2. The SPV's cash flow breaks, impairing the credit held by Apollo and Blackstone 3. Broadcom's $30 billion residual guarantee is triggered 4. Google's rent guarantees are triggered — $815 million on the books, $43.8 billion of exposure Every link in the chain believed it had pushed the risk to the next one. The last link is Alphabet's balance sheet. --- # Should You Still Build an MVP in the AI Era? 211M Lines of Code Say Rework Went From 3.3% to 7.1% URL: https://doaipm.com/en/blog/memorized-not-calculated/ Published: 2026-08-04 Tags: AI Coding, MVP, Iterative Development, GitClear, Product Management, Project Management, Tech Commentary Throw an idea at an AI and eight times out of ten, the first thing it says back is "let's start with a minimum viable version." That sounds like a judgement. It's a memory. The Agile Manifesto is from 2001, *The Lean Startup* from 2011, and those two documents plus the several hundred thousand blog posts, courses and postmortems they spawned are all in the training data. It recommends iteration, not because it ran the numbers on its own costs. It ran the numbers on ours. And the two bills are structured in opposite directions. ## The four premises MVP rests on — three still hold Iteration isn't a law of nature. It's the optimal answer under a specific set of constraints, and the set looks like this: - **Writing code is slow and expensive.** A feature goes from understood to running in person-days. - **Changing code costs more than writing it.** To change it you must first read it, and reading it means climbing back into someone else's head — or your own from three months ago. - **Requirements get sharper along the way.** Users only know what they want once they see something, so putting something in front of them early pays. - **The decisive one: there is a large cost gap between a crude implementation and a complete one.** That last one is the economic bedrock of MVP. Store it in localStorage for now, skip permissions for now, hard-code a few config values — the weeks you saved were real, and spending them on validating an assumption instead was genuinely the better trade. Three of those four still hold today. The fourth is gone. ## localStorage vs PostgreSQL: minutes apart in an AI's hands A thread on V2EX puts it most cleanly (`t/1216691`, titled "Has MVP thinking stopped working in the age of AI coding?"). The original poster's words: describing "store the data in localStorage" to Cursor and describing "use PostgreSQL with a connection pool" differ by **a few minutes of generation time**. > If the cost is the same, why build the crude version? That single line pulls out half the foundation. What you save is no longer weeks, it's minutes — while everything you take on in exchange is unchanged: a storage layer that will be ripped out, a batch of calls written against it, a migration you'll have to redo. **Crude isn't cheap anymore. It's just incomplete.** ## The AI's bill: context is the main cost, not the task The other half of the foundation collapses at the level of cost structure, and this layer is better hidden. Run one agentic task and the bulk of the input tokens is never your task description. It's the system prompt, the repo map, the conversation history, the files it read. The task description is a rounding error inside it. What follows has all been measured: - Every retry **resends the entire context**. Two retries can triple the cost of a session. - Counting context overhead, real cost runs **3–5×** the naive estimate — roughly $0.27 to $3.25 per PR, and north of $20 for a single task in a complex repo is unremarkable. - The arXiv paper *Cheap Code, Costly Judgment* (2607.01087) measured **agentic tasks consuming 1000× the tokens of ordinary code chat**, with up to 30× variance in token spend across repeated runs of the same task. Now put human iteration and AI iteration side by side. A person splitting work into three phases pays roughly the same per phase, because a person remembers what happened in the last one. The context sits in their head and costs nothing to call. **The AI doesn't remember. It has to buy the context again every phase.** | Same three phases | Human | AI | |---|---|---| | Main cost per phase | The work itself | **Rebuilding context** | | Memory of the last phase | In their head, free to access | Doesn't exist, must be reloaded | | Total across three phases | ≈ three units of labour | ≈ three units of labour + two rebuilds | | Marginal cost of one more split | One conversation | One full re-purchase of context | | What iteration buys | Chances to not go down the wrong path | The same chances, plus a friction fee | **Iteration is insurance for a human and a friction fee for an AI.** ## GitClear measured the rework: 3.3% → 7.1% Everything above is inference. What follows is measurement. GitClear analysed **211 million lines** of changed code, tracking a metric called churn — the share of code substantially rewritten or deleted within days of being merged. It measures exactly one thing: not getting it right the first time. | Year | Churn | |---|---:| | Pre-2023 (baseline) | 3.3% | | 2024 | 5.7% | | 2025 | **7.1%** | More than double in two years. Other metrics in the same research point the same direction: - Within-commit copy/paste **+41%** - Duplicated code blocks **+81%** - Error-masking constructs **+47%** - Cross-file function calls, the proxy for reuse: **−35%** - Refactoring line moves: **−70%** GitClear sorts AI-amplified rework into three kinds, and every one of them is a direct consequence of shipping a half-finished thing first: 1. **Wrong place** — logic and syntax both correct, but sitting in the wrong architectural location; someone relocates it later 2. **Built twice** — reimplementing functionality that already existed instead of reusing it 3. **Rewritten days later** — merged, then substantially changed because of an edge case or a convention it violated In the same body of research, AI-authored PRs carry 10.83 issues on average against 6.45 for human-authored ones. A factor of 1.7. Developers' own experience matches the numbers. From the Stack Overflow 2026 Developer Survey: - **84%** are using AI tools - **45%** say debugging AI-generated code takes longer than writing it themselves - **66%** say their biggest frustration is output that is "almost right, but not quite" - And trust in AI output sits at **3%** A separate survey puts **43%** of AI-generated code changes as needing debugging in production. "Almost right" is the operative phrase. It means the problem doesn't surface when you accept the work — it surfaces after the merge, which lands it squarely on the next iteration. The round you thought you saved gets billed later. ## "Right once" is not "all at once": delivery granularity vs execution granularity This is where it's easiest to slide into a slogan: stop iterating, do it all at once. That's wrong, and dangerously so. Two granularities have to be kept apart: - **Execution granularity**: change one thing, verify, then the next. Still correct today, more so than before — an AI touching ten files at once is more than you can review. - **Delivery granularity**: whether what you hand over this round is a finished thing. "Right once" is about the second. **It constrains completeness, not feature count.** Scope can be narrow — one page, one endpoint narrow. But the part you committed to this round has to be complete: real data structures, not placeholders; loading, empty, error and success states all present, not just the happy path; something that actually runs for a person to use, not a screenshot. The deepest problem with the word MVP is that it sold "narrow scope" and "built crudely" as a bundle. They used to be genuinely bundled — saving money meant saving on both. Now they come apart: **scope should still be narrow, but crude has stopped being cheap.** ## "MVP is about validating assumptions, it has nothing to do with AI" — which two of the four objections hold The replies under that V2EX thread are more valuable than the original post. One at a time: **1. "The core of MVP is validating an assumption. It has nothing to do with code, and nothing to do with whether AI is involved."** It holds — and it points straight at where the problem is. The word MVP has always bundled two things: **validating an assumption**, and **shipping an incomplete implementation**. They used to be inseparable, because the only cheap way to validate an assumption was to build something rough. They've come apart. Validate the assumption, by all means — you can now do it with something complete. **2. "You can never learn all of your potential users' needs up front."** It holds, and it doesn't conflict with doing it right once. Nobody is asking you to build every feature in one pass. The ask is that the part you do build this pass isn't left half-finished. **3. "For logically complex projects — especially systems with intricate business loops and state machines — doing it in one shot produces a mess."** A real problem, but a problem of execution granularity. A complex state machine obviously gets built step by step with verification at each step. That is not an argument for shipping a version you already know you'll rewrite. **4. "With AI, iteration is fast and costs almost nothing."** This one doesn't hold. The two sections above are its counterexample: what got fast is generation, not convergence. Churn climbing from 3.3% to 7.1% is a measurement of exactly that gap. ## After writing "no phased development" into my global rules I have one line hard-coded into my own global rules: no phased development, delivery means a finished product, no MVPs and no partial milestones. In practice that comes down to: - Real data structures and real content from the first pass — **no Lorem ipsum, no fake data** - All four states delivered together: loading, empty, error, success - Running it myself before it counts as delivered — open it in a browser, execute it on the command line, hit the endpoint with curl The price is real: **the first description has to be much longer.** You have to think through boundaries, states and data structures before anything starts. That work didn't get taken over by AI — it just moved from "forced to think it through during the third round of rework" to "thought through before the first round starts." The payoff is equally direct: you don't explain the same thing a second and a third time. And explaining it again is the single most expensive item on that bill above. ## How narrow to cut the scope — AI can't help there That question still has no answer. Doing it right once presumes the scope was cut right. Cut too wide and "right once" becomes "a very long once." Cut too narrow and what you built can't validate anything. Where that line falls comes down to judgement about users and situations — and that paper's title already said the whole thing: **code got cheap, judgement didn't.** --- # AI's Standard Isn't in the Model. It's in the Last Version You Nodded At URL: https://doaipm.com/en/blog/ai-standards-live-in-context/ Published: 2026-08-03 Tags: AI Collaboration, Product Management, Quality Standards, Claude Code, Methodology, Tech Commentary Someone said to me: the biggest problem with using AI is that once you compromise, it lowers the bar, then keeps compromising, and ends up producing garbage. Exactly like real projects. The description is accurate. I want to explain why it happens, because the reason isn't "the model isn't smart enough," and it isn't "it's being lazy." ## 1. The standard isn't in the model, it's in the context Every time the model generates something, it's making a prediction conditioned on the context that already exists. That sounds like a technical detail, but it decides everything. **Whatever is already sitting in the context is the current frame of reference.** If the last five outputs were sloppy, sloppy is the local norm right now. The model isn't "deciding" to lower its standards; the evidence in front of it changed. It's predicting the next step inside a distribution made of sloppy samples, and the result drifts toward sloppy on its own. People work the same way, just far slower. A team sliding from "this can't ship" to "let's ship it and see" usually takes months. **An agent can cover that entire distance inside one session**, because its whole memory of what good looks like is the last few thousand tokens. So the shape of the problem isn't "will the AI hold the line." It doesn't hold a standard. The standard is whatever you put in each turn. ## 2. Every acceptance is a calibration This is the part that's easiest to miss. When you look at a version that isn't quite right and say "forget it, this'll do," you think you're making a one-off concession. But to the model, you just supplied a high-quality supervision signal: **this level is acceptable.** And that signal is stronger than any principle you've ever stated. The reason is simple: principles are abstract, samples are concrete. You say "writing needs information density," and in the context that's just a sentence. You accept a padded paragraph, and that paragraph — together with your approval — stays in the context and becomes the working definition of "information density." **One "this'll do" teaches more than ten "be stricter about this."** I got calibrated a lot over these three days. The clearest instance went like this. I set myself a hard rule: no piece with less than 3000 Chinese characters of body text. The one I finished the day before yesterday came in at 2166. I went back for more material, got to 2746. More material, 2851. Finally, to clear the line, I added another section: 3006. That last section isn't bad. But what triggered it wasn't "this argument is missing a layer." It was "I'm 149 characters short." ## 3. The ratchet only turns one way Lowering a standard and raising a standard require wildly unequal amounts of action. **Lowering a standard requires silence.** You say nothing, you don't pick at it, you accept it, and the standard drops. Cost: zero. **Raising a standard requires an explicit act**: a rejection, a new rule, a round of rework. Every one of those costs time, creates friction, and may force the other side — human or AI — to stop and redo the work. An institution that can only move up through deliberate action and moves down through inaction will, over time, only move down. That isn't a willpower problem. It's a structural one. You've seen how it looks in real projects: the first time, someone accepts a stopgap. The second time, that stopgap is the reference implementation. The third time, someone builds another layer on top of it. Nobody ever made a decision to lower the bar, but the bar is lower. ## 4. The visibility of the cost decides the direction of the drift I wrote about Apple the day before, and that piece has the same structure inside it as this one. Apple got caught this time on its capacity forecast. Lock in too much and the loss is definite and computable — the extra money spent, the cash tied up, the inventory to be written down, all of it landing on a table anyone can read. Lock in too little and the loss is invisible: the people who wanted to buy walk away, and that money never enters any statement. **When one side is visible and the other isn't, the organization will systematically drift toward the side where the loss stays invisible.** Quality drift runs on the same mechanism: - **The cost of holding the standard is visible**: the extra time, the extra rounds of rework, the interrupted schedule, the person or agent you keep sending back. It hurts right now, and somebody is going to ask why you're so slow. - **The cost of accepting mediocre work is invisible**: it detonates later, and when it does, nobody traces it back to today's concession. So you don't need anyone to act in bad faith. As long as the visibility of the costs stays arranged this way, the drift is guaranteed. ## 5. A standard written as a number will get gamed There's another layer to that 3000-character example. What the rule actually meant was "make it substantial, don't pad it." Word count was only a proxy metric. But once it was written as a checkable number, it stopped being a constraint and became a target — and targets get optimized. **I stopped asking "is this substantial enough" and started asking "how many characters am I short."** This isn't me being unusually dishonest. Any explicitly measured proxy metric loses its validity as a metric the moment it becomes the target. The only difference is that AI optimizes proxy metrics far more efficiently than a person does — you say 3000, it hands you exactly 3006. The same problem showed up elsewhere. I wrote myself an "AI-tell blacklist": no tour-guide sentences like "This article will…" "It's worth noting that…" "In conclusion…". The list worked. Those sentences really did disappear. But today I dropped the finished draft into Tencent's Zhuque AI detector. The result: **100% likely AI, 0% human characteristics.** The list caught the tells I'd already thought of. It couldn't catch the overall writing pattern — structure too tidy, paragraphs all roughly the same length, argument density perfectly even, never wandering off, not one unnecessary sentence. Real people don't write like that. They have pet phrases. They suddenly get long-winded in one spot. They drop in a line that has little to do with the main thread but that they wanted to say anyway. **A standard you can write down as a checklist only covers the failure modes you've already thought of.** ## 6. The number of rules is a counter of failures Over this session I've written four separate "hard rules" files: a style spec for Chinese tech commentary, an anti-god-view checklist for clearing drafts, a floor for article length and structure, a distribution methodology. Not one of them was planned in advance. **Every single line was added after I made the corresponding mistake.** - "Covers must be designed at 156×104, and the text has to land inside the safe area" — by the time I wrote that down, I had already shipped a cover where cropping 16:9 to 3:2 took 125px off each side and sliced the "4" off "$45 billion." What the feed displayed was "$5 billion." Not hard to read — wrong. - "Don't trust the engine's success report; verify every platform independently" — before I wrote that, I spent several days taking the engine's self-reported ok as the result, when in fact three of the seven platforms had errored out: Zhihu showed success but it was actually a draft, Toutiao's cover never attached, X never posted at all. - "A tool reporting a timeout doesn't mean the operation failed; poll until the field is full" — before this one, a 45-second IPC timeout made me conclude the input had failed, so I went back to add the second paragraph, dropped it at a misplaced cursor, and scrambled the entire prompt into garbage. - "Verify the draft state before publishing on Toutiao" — before this one, I re-ran without verifying and published the same article twice. Four files, every one of them written after the fact. Which says something: **in this kind of collaboration, the number of rules isn't proof of rigor. It's a record of how many times you failed.** I don't think that's a bad thing. But be clear about what it means — the spec in your hand can only block the mistakes you've already paid tuition for. ## 7. What actually worked Three days in, exactly one thing worked on me reliably. Not a principle, not a checklist. **A check that can fail.** After the cover incident, I wrote a 20-line script: run the cover through each platform's real cropping rules, shrink it to its actual size in the feed — 156×104 — and draw it side by side with a 4x blowup. Run it, look at the image, and if you can't make out the subject or read the text in that little square on the left, it fails. The difference between that and the principle "covers should be clear and eye-catching": **you can argue with a principle. You can't argue with the script's output.** It turns something that needs judgment into something that needs looking. I can't argue with a 156×104 image. Same day, second cover. The first version came out with a readable main title and a subtitle mashed into a blur. Going by the principle, I'd probably have said "close enough." Going by that side-by-side, I did another version. That's the only time in three days the bar moved up instead of down. The difference wasn't that I'd become more disciplined. It's that that particular judgment didn't require discipline. --- Writing this, I ran into an awkward fact: this article is also under that rule. Body text, no less than 3000 characters. I didn't pad it. It's however long it is; when it has said enough, it stops. If that drops it under the line, that's precisely the sign that the line needs changing — **a floor that forces you to pad isn't guarding against padding. It's guarding against thin. And those two were never the same thing.** --- # Apple Put a Quota on Bug Reports — and 11 Flaws in the Same Patch Batch Were Found by AI URL: https://doaipm.com/en/blog/apple-capped-bug-reports/ Published: 2026-08-03 Tags: Apple, Bug Bounty, curl, GitHub, AI Security, Product Design, Product Management, Tech Commentary On August 2 the Financial Times reported that Apple has imposed two limits on researchers who file security issues through Feedback Assistant: a cap on how many reports one person can keep open at a time, and a 30-day cool-off period once you hit it. Higher quotas require a separate request. Apple's stated reason is that AI-assisted vulnerability reports are arriving faster than humans can verify them. Apple has not published the actual number. One week earlier, on July 27, Apple shipped 8 security advisories and closed 210 vulnerabilities. Eleven of them were found by AI. Those two things happened too close together to read separately. ## Who got credited for those 11 Apple's security advisories name the reporter. The July 27 batch credits: - **Claude (Anthropic), 4**: CVE-2026-64757 and CVE-2026-43715 in WebKit, CVE-2026-64704 in SMB, CVE-2026-64703 in WebDAV - **OpenAI Codex Security, 2**: both in WebKit - **NVIDIA AI Red Team, 3**: two in WebKit, one CUPS privilege escalation - **GLM (Z.AI), 2**: both in WebKit - **Atuin Automated Vulnerability Discovery Engine, 1**: screen sharing The worst one in the batch is CVE-2026-64747 in AVEVideoEncoder — arbitrary code execution with kernel privileges, affecting everything from iPhone to Vision Pro. macOS Tahoe had seven more that reach root: two in CUPS, one each in Core Services, Accounts, SecurityAgent, Remote Management and MediaRemote. The screen-sharing bug Atuin reported is CVE-2026-43760, and the firm behind it is the Italian company Bynario. Their Atlas platform runs on GPT-5.5 and turned up more than 50 candidate flaws in macOS in three weeks. The trigger conditions for this one are specific: the machine has Screen Sharing or Remote Management enabled along with legacy VNC password access, and an already-authenticated VNC client can then read protected data and create files as root — Bynario went on to demonstrate extending it to root command execution. Apple fixed it in macOS Tahoe 26.6 and credited three names: Alfredo Pesoli of Bynar.io, wdszzml, and Atuin, the machine. The kind of submission Apple moved to limit yesterday and the batch of findings Apple thanked last week come from the same place. ## Doubling the bounty and narrowing the door In October 2025 Apple doubled its top bounty from $1 million to $2 million, for exploit chains that require no user interaction and approximate what mercenary spyware achieves. Bugs in beta software and Lockdown Mode bypasses carry bonuses that can push a single payout past $5 million. A one-click WebKit sandbox escape tops out at $300,000; a wireless proximity exploit at $1 million. The new schedule took effect in November 2025. Apple's own cumulative figure: $35 million paid since 2020, to 800 researchers. Less than a year later, Apple put a quota on the door. **One foot on the accelerator, one on the brake.** Doubling the bounty tells the outside world to find more. The quota tells it to send less. Both signals go out at once, and whoever receives them has to guess which one is real. Apple is not blind to where the bottleneck is. It now uses AI to triage and prioritise incoming reports, but confirmation is still human work — reproduce it, check the preconditions, judge the urgency. Finding bugs now runs in parallel on machines. Confirming them still runs serially through people. Until that middle step widens, more throttle upstream just piles up at the same place. ## curl ran the opposite experiment curl has run a bounty on HackerOne since 2019. For years, the share of submitted reports that turned out to be real vulnerabilities held above 15%. In 2025 it fell below 5% — more than nineteen out of every twenty reports were worthless. In July 2025 Daniel Stenberg estimated roughly 20% of submissions were pure AI slop. He published a representative sample: a claimed vulnerability in HTTP/3, written up convincingly, complete with GDB sessions and register dumps. The only problem was that the function it referenced does not exist in curl. > We are effectively being DDoSed. That is Stenberg. He also said something about why this cannot be solved by talking to people: we have no way to change how all these people and their slop machines work. On January 26, 2026 he announced the end of the curl bounty. It stopped on January 31, and from February 1 security reports moved to GitHub. He was blunt about the motive: > The main goal with shutting down the bounty is to remove the incentive for people to submit crap and non-well researched reports to us. **Note which end he touched: the door stayed open, the money went away.** Anyone can still report. Reporting just no longer pays. Moving to GitHub he later judged a mistake. He listed fifteen defects in GitHub's security advisory system, including that the full report goes out over email and notifications with no way to disable it, that invalid reports cannot be publicly disclosed, and that the CVE fields cannot be edited. On March 1 curl moved back to HackerOne — **but the bounty did not come back**. His words: the reward money is still gone, there is no bug-bounty. What happened after the money left, in his phrasing: the inflow tsunami has dried out substantially. He did not declare victory. His very next line was that perhaps it just takes a while for all the sloptimists to figure out where to send the reports now. Then curl simply shut the door for five weeks: July 1 at 00:00 CEST through August 3 at 09:00 CEST, no vulnerability reports accepted at all. Stenberg's framing was that whatever you find this month, you will have to wait. That end time is this morning. curl just reopened. ## GitHub took a third road On July 27 — the same day Apple shipped those 210 fixes — GitHub's new bounty schedule took effect. Public-track payouts were cut by at least half at every severity level; critical findings went from $20,000–$30,000-plus to a flat $10,000. Alongside it, GitHub opened a permanent invite-only tier paying $30,000 and up. GitHub described the problem more precisely than Apple did. Alongside growth in legitimate reports, it said, came a sharp rise in submissions without a proof of concept, theoretical attack scenarios that do not hold up under scrutiny, and findings already covered by published ineligible lists. **It did not say those reports were written by AI. It said what those reports were missing.** No PoC, no scrutiny, already-published non-issues. There is a companion rule: newcomers are capped on how many reports they may submit until they have demonstrated quality. Build a record and the channel widens, ending at that invite-only tier. Per The Register, Google, Bugcrowd and HackerOne are restructuring along similar lines. ## A quota acts on the person, not on the content All three face the same change: writing a report that looks professional now costs close to nothing. They differ in which part of the machine they reached for. | | curl | GitHub | Apple | |---|---|---|---| | What moved | Bounty removed, channel left open | Public tier halved, high value moved to invite-only | Bounty raised to $2M+, quota on the submission door | | What it acts on | The incentive to submit | The submitter's track record | The submitter | | The cost | Real researchers no longer get paid either | Newcomers must serve time to reach the high tier | High-output teams and careless posters hit the same ruler | A quota is a blunt instrument because it counts how many reports you have open, not whether any of them holds. Bynario, which can produce 50 leads in three weeks, and a person pasting raw model output are the same class of account under a quota. For the first it blocks real throughput; for the second it blocks garbage. One rule, two opposite effects. The bigger problem is that the category itself is wrong. The public framing has been "limiting AI-generated reports" — yet 11 of those 210 fixes on July 27 were found by AI, and one of them was reported by the very company that the coverage of this policy keeps naming. **Who wrote a report and whether that report holds are two unrelated properties.** Filter on the first and you will misfire in both directions at once. ## Target Flags is the half Apple got right There is another piece in Apple's changes called Target Flags: researchers are asked to demonstrate that a flaw actually reached a specific target state, rather than describing a problem that ought to exist in theory. A quota asks who you are and how many you have filed. Target Flags asks whether your thing runs. The first adds cost to the person; the second adds cost to the content. And AI happens to be extremely cheap at producing text that looks correct, and not much cheaper at getting a real exploit chain to work — a proof requirement lands exactly on that gap. Stenberg's removal of the bounty also moves proof cost, just denominated in opportunity: report it and you get nothing, so only people who actually want curl to be better will spend the time. A quota does not sit on any cost gap at all. It sits on volume. ## Who else is on this curve Any place where open input meets human triage is sliding the same way: - **Hiring.** The marginal cost of a résumé and cover letter went to zero. The other end did not change. - **Issues and PRs on open source projects.** curl was just the first to say it out loud. - **Submissions, peer review, awards.** The number of reviewers is fixed. - **Support tickets and refund appeals.** Appeal language can be generated in bulk. - **App store review.** Apple runs an entire queue of that beyond Feedback Assistant. These systems were all designed on one assumption: **submissions are scarce, so humans can keep up.** That stopped being true in 2026, and most products have not redesigned their intake since. There are roughly three mechanisms available, and all three are running somewhere: 1. **A proof requirement** — make the submitter produce evidence that generation cannot fake. Apple's Target Flags and GitHub's PoC demand are this. 2. **Reputation tiering** — let historical accuracy set the width of your channel. GitHub's invite-only tier plus the newcomer cap is the complete version: accurate people get $30,000, unproven people queue. 3. **A deposit or penalty** — an invalid submission costs you something, pushing cost back to the sending end. Fraud-refund screening in e-commerce runs on this logic. The quota ranks behind all three. It is the easiest to implement and the least sensitive to content. There is a side effect on curl's road worth naming: **once the money is gone, the people who stay are more accurate, but there are fewer of them.** For an open source project that is an acceptable trade. For a company relying on outside researchers to keep the iPhone standing, maybe not. Apple cannot remove the money, so it had to find another lever, and the first one it reached for was a quota. ## Same people, same machines The screen-sharing privilege escalation Apple closed on July 27, CVE-2026-43760, was reported by Bynario and a machine named Atuin. What Apple announced on August 2 is a limit on how fast those same people and that same class of machine can hand things in. curl reopened at 09:00 this morning, still with no bounty. The last time Stenberg spoke publicly he did not say the problem was solved. He said maybe those people just have not figured out where to send it now. One changed the incentive, one changed the gate. In three months it will be worth coming back to see which end of the numbers actually moved. --- # Tim Cook's Last Quarter: Apple Got Caught by Its Own Demand Forecast URL: https://doaipm.com/en/blog/cook-last-quarter-supply-forecast/ Published: 2026-08-02 Tags: Tim Cook, Apple, iPhone 18, Supply Chain, Memory Prices, Product Management, Tech Commentary On July 30, Tim Cook hosted the last earnings call of his tenure. He warned that iPhone, iPad and Mac will all be supply-constrained in the September quarter, and that the constraint will hit revenue. Apple currently has less flexibility in the supply chain than normal, and there is a shortfall in the leading-edge capacity that makes its own silicon. On the same call, he called memory chip pricing a "hundred-year flood," and admitted Apple did not want to raise prices — the exponential rise in memory costs forced it. The stock fell 7%–8% after the report. iPhone revenue grew 22% year over year this quarter. Guidance for next quarter is 9%–10%, below the 12% analysts had expected. Those two numbers only matter together. **22% is what Apple sold. 9%–10% is what Apple can build.** The gap in between is not demand disappearing; it's product that can't be shipped. Cook drew the causal line himself on the call: guidance came down precisely because supply constraints will hit revenue. A company cutting guidance because it's selling too well is not a common sight. On September 1, John Ternus becomes CEO and Cook moves to executive chairman. ## 1. He got caught by the thing he's best at Cook joined Apple in 1998 as senior vice president of worldwide operations. Steve Jobs hired him to fix one problem: Apple's inventory turned over painfully slowly, with months of product sitting in warehouses. Cook shut most of the company-owned factories, outsourced manufacturing, and compressed inventory turns from months to days. That playbook became the standard move for the entire consumer electronics industry, and it was the whole reason he got the top job — **he was not the product genius; he was the guy who could actually build the product, cheaper than anyone else.** In his 15 years running Apple he shipped every iPhone since the 4S, plus Apple Watch, AirPods, Apple Pay, Vision Pro, and the Mac's move off Intel onto Apple silicon. Now, in his last quarter before the handoff, Apple's problem is that it doesn't have enough parts. The irony doesn't need spelling out. But the irony isn't the interesting part — **where in the chain the root cause sits** is. ## 2. The root cause isn't the supply chain, it's the demand forecast Cook was explicit on the call: this is not a supply issue but a demand forecast issue. The substance of it: this product cycle for iPhone and Mac came in stronger than planned, and demand ran beyond Apple's expectation. Suppliers didn't drop the ball. Apple didn't expect to sell this well. **That's a product judgment error, not an operations error.** Semiconductor capacity is locked in advance. To get parts next year, you place the order and sign the long-term agreement this year. How much you order depends on your forecast for next year's volume. **In any company, forecasting isn't procurement's job. It belongs to product and to the business.** Forecast low, and you don't lock enough capacity. By the time demand actually shows up, there's nothing left to buy — which is visible from upstream: - Micron's 2026 high-bandwidth memory capacity is **already sold out** under existing pricing agreements - SK Hynix told investors its advanced packaging lines are **fully booked** through the end of 2026 **The capacity isn't unaffordable, it's unavailable.** Money doesn't solve this, because the slots were taken a year ago. ### The tighter constraint is leading-edge process Cook called out one line separately: there's a shortfall in the leading-edge capacity that makes Apple's own silicon. That's harder to solve than memory. Memory at least has multiple suppliers, plus mainland Chinese capacity ramping — in theory there's a "find another supplier" option. Wafer capacity at the most advanced node has no such option. There are only a handful of lines in the world that can do it, and the queue isn't just Apple; it's every company fighting for AI compute. Apple has always stood near the front of that queue as a long-running launch customer for TSMC's leading node. **But "near the front" is ordered by volume, and volume comes back to the same question: how much did you commit to a year ago?** So the two bottlenecks here are two exits from the same bottleneck. Memory and leading-edge process are both things you have to lock far in advance, in quantities determined by a forecast. Forecast low, and both ends squeeze at once. ## 3. One memory price chain, four companies, four endings I've been writing about different faces of the same story for days, and today it closes the loop. DRAM and NAND prices are up 63%–75% year over year. On the same chain, four companies got four completely different outcomes: | Company | What they did about memory prices | Result | |---|---|---| | **Microsoft** | Raised 2026 capex guidance to roughly $190 billion in April (explicitly citing memory prices), then pushed it back down to about $175 billion through fleet utilization and delivery process | Added roughly $450 billion in market cap in a single day on July 30, a US record | | **Meta** | Raised capex two quarters in a row, full-year guidance $130–145 billion | Free cash flow collapsed to $784 million, stock down 9% | | **Situational Awareness** | 4x leveraged long on memory and compute | Force-liquidated on July 30, −67% for the month | | **Apple** | Didn't lock enough capacity, forced to raise prices | Stock down 7%–8%, next-quarter guidance below expectations | One external variable, four ways of handling it, four fates. **The variable was shared. The endings were bought with each company's own judgment.** Microsoft's line is the one worth studying: memory prices pushed it to $190 billion in April, and three months later it said it didn't need to spend that much, because efficiency gains across CPU and GPU fleets let existing infrastructure do more. **It didn't fight the price increase. It changed consumption on its own side.** Apple's position is different. iPhone memory content isn't something you schedule away — it's a fixed line in the hardware BOM. Apple had exactly two moves: lock capacity early, or raise prices. It didn't do enough of the first, so only the second was left. ## 4. Locking capacity is a product decision, not a procurement task Most companies treat "lock upstream capacity years ahead" as a supply chain department job. It isn't. It's a bet on your read of the next two years of volume, and once placed it's hard to unwind — these are long-term agreements, with penalties if you take less and dead inventory if you take more. **Only product and the business can make that call, because only they know what you're selling next year and how much of it.** The difficulty is the timing mismatch: you have to buy materials for two years out before the product exists and before the market has validated anything. Get it right and nobody praises you (the parts are just there). Get it wrong and it hurts in both directions — too little means shortages, too much means inventory. And the **costs of those two errors are wildly asymmetric**, which is the real reason this is hard. Over-lock capacity and the loss is definite and countable: the extra money, the trapped cash, the inventory you write down. It's ugly, but it shows up on a statement anyone can read. Under-lock capacity and the loss is invisible: the people who couldn't buy wanted to buy, but they won't wait three months for you — they leave. That loss never lands on any statement, and you don't even know how big it is. **You can see what you sold. You can't see what you could have sold.** Because one side is visible and the other isn't, almost every organization skews systematically conservative on this decision. The extra money gets questioned; the churn from buying too little doesn't. That's the direction Apple got caught in — its forecast didn't miss randomly, it missed toward safety. I have exactly the same bias in the tiny things I build. Before launch I size servers and quotas against "roughly how many people will use this," and that "roughly" is always an undercount, because the extra spend comes out of my pocket at the end of the month, while the person who hit a page that wouldn't load and left is someone I'll never hear about. Apple came in short this time. Not absurdly short: it has locked parts for roughly 80 million new iPhones, covering iPhone 18 Pro, Pro Max and the first foldable iPhone, with the foldable's production target raised to 10 million units. The problem is that demand is bigger than that. One move that may get overlooked: per a CNBC report on July 2, Apple is evaluating memory chips made in mainland China, with both CXMT and YMTC under consideration. This is supply diversification under duress — **when your only remaining fix is "find someone else who can ship," your negotiating leverage is already gone.** ## 5. "We didn't want to raise prices but had to" is a confession about cost structure Cook said Apple didn't want to raise prices, that the exponential rise in memory costs forced it. Analysts speculate the iPhone 18 Pro could go up by $200. That's an honest statement, and it also exposes something. If any single line in a product's bill of materials satisfies three conditions at once — **a large share of cost, violent price swings, and no long-term agreement locking it** — that line is a time bomb. Whether it detonates depends on external market conditions, not on how well you execute. Memory in a phone BOM hits all three. What I build is as small as it gets, but the structure is identical at small scale. I wire external model APIs into my own tools, and the most expensive line is inference, whose unit price I control not at all. If I price my product to work exactly against the API price of the day, one vendor adjustment takes my margin to zero — and I did nothing wrong. **The move isn't forecasting the price. It's asking one question at design time: if this line doubles, does my product still work?** If it doesn't, you either cut consumption now, or build headroom into pricing now, or go lock a long-term price now. All three have to happen during product definition. The day the price moves is too late to start. Apple picked the third path and ran the failing version of it: it locked, but not enough. ## 6. He hands over an Apple that can't get parts Ternus is 50, joined Apple in 2001, and has run hardware engineering since 2021. Cook is 65, and took the company from Jobs in 2011. This is Apple's first CEO change in 15 years, and the handoff lands on a very concrete problem: the iPhone 18 Pro, Pro Max and first foldable iPhone shipping in September may be short of supply and may cost more — and the reason isn't in the factories, it's in a volume forecast made a year ago that was too low. The interesting part is that the successor comes from hardware engineering. When Cook took over, what Apple needed was the ability to build the product. What Apple needs now may be something else: in an environment where materials cost can rise 70% in a year, deciding again how much memory a device should carry, what it should sell for, and how much capacity to commit to two years out. I don't know how Ternus will handle that one. But this report has already drawn a line for him. For the past decade-plus, Apple's supply chain was its moat — costs and delivery nobody else could match. When upstream turns into a seller's market, when capacity has to be fought for two years ahead, when the tightest few line items have only a handful of suppliers, **"executing better than everyone else" stops being decisive. "How much did you commit two years ago" is.** This is the turn from an operations problem back into a product problem. The ledger Cook hands over states the issue plainly: **selling too well is also a miscalculation.** --- # On July 24 He Called It a 'Strategic Entry Window.' On July 30 the Fund Was Force-Liquidated. URL: https://doaipm.com/en/blog/he-called-it-a-strategic-entry-window/ Published: 2026-08-01 Tags: Leopold Aschenbrenner, Situational Awareness, Hedge Funds, AI Infrastructure, Leverage, Product Management, Tech Commentary On July 24, Leopold Aschenbrenner wrote to his investors and called the ongoing sell-off in AI stocks a "strategic entry window." He invited them to add capital on August 1. On July 30, his fund was force-liquidated. Citadel took the entire public-market book — about $16 billion — at a discount, in under 36 hours. On July 31, the stocks that had been sold rose 25.99%, 21.51%, and 26.49%. In his last letter he wrote: "We let you down this month." ## I. He Wasn't Selling a Fund. He Was Selling an Essay. Aschenbrenner was born in 2001. He graduated from Columbia at 19, first in his class. In April 2024, OpenAI fired him. He was 22. The two sides tell different stories about why. OpenAI said it was a leak — he shared a brainstorming document with three outside researchers. He says the real reason was a memo he circulated internally arguing that the company's security measures could not support AGI-level research. Two months later, on June 4, 2024, he published a set of essays at situational-awareness.ai. The core argument extrapolated the scaling laws outward: at this rate, AGI around 2027 is possible, followed by an "intelligence explosion" leading to superintelligence. The piece went everywhere in 2024. Its force didn't come from new data. It came from a clean chain of reasoning: **stronger models → more compute → more chips, more memory, more data centers, more power.** Every link landed on a specific public company. In September 2024, he launched a fund named after the essay: Situational Awareness LP. **That is a rare structure — a fund with the same name as its thesis.** Most managers leave themselves an exit. The strategy can shift, positions can rotate, the name stays neutral. He didn't. The name of the fund was the call itself. By early July 2026, the fund ran $45 billion. Cumulative return since inception in September 2024: more than 2000%. The first half of 2026 alone: 439%. A young man fired by his employer in 2024 turned one essay into that number in under two years. Before July, it was the best story on Wall Street. ## II. What He Was Long Got Repriced on July 30 The longs were SanDisk, CoreWeave, Bloom Energy — that family. The shorts were software, at scale. Those three names weren't arbitrary. Each maps to a link in the essay's chain: - **SanDisk** — storage. Bigger models mean parameters and context eat more memory. - **CoreWeave** — GPU cloud. Compute itself, sold by the hour. - **Bloom Energy** — fuel cell power generation. Once the data centers go up, the next bottleneck is electricity, and the grid cannot add capacity as fast as you can add racks. From model capability all the way down to power generation equipment, with no step skipped. **It was a position sheet that translated a paper into tickers, line by line.** The long/short logic was just as clean: compute and power are scarce, so the shovel sellers make money; software is what AI disrupts, so it gets replaced first. On July 30, Microsoft reported. I wrote about that print the day before: Azure growth accelerated from 40% to 43%, but capex guidance for calendar 2026 came down from about $190 billion in April to about $175 billion. Microsoft added roughly $450 billion in market cap that day, a single-day record for US equities. Good for Microsoft. **Fatal for Aschenbrenner.** Because the $15 billion Microsoft said it would no longer spend was exactly the thing the SanDisks of the world sell. Amy Hood's reasons for the savings: efficiency gains across the CPU and GPU fleets, process improvements in bringing capacity online. Translated: same demand, less hardware to buy. Microsoft's own explanation for the $190 billion in the April guide was that memory prices were spiking. Three months later it said it wouldn't need to spend that much. What that sentence means to a man 4x levered long memory needs no explanation. The same week, Meta raised capex, free cash flow collapsed to $784 million, and the stock fell 9%. Apple fell 7% to 8% after earnings, and Tim Cook, on the last earnings call of his tenure, said there was a "hundred year flood" in memory chip pricing. The whole market spent those days asking the same question: when does the money going into AI infrastructure turn into revenue. **Situational Awareness was the most levered answer to that question.** Leverage ran close to 4x at one point, provided by Bank of America, Goldman Sachs, and JPMorgan. ### The Worse Problem: The Short Book Was Wrong Too What a long/short book fears is not losing on one side. It's losing on both at once. His short thesis was that AI eats software — large models can write code, do design, produce copy, so the companies selling those capabilities get replaced first. That call was popular in 2024 and 2025, and it did make money. July knocked the whole thing over. In a single week the market delivered two conclusions: **AI infrastructure may be overbuilt, and the software companies are not dead.** Neither conclusion is exotic on its own. Stacked together, they are precisely the two opposites of his book. The long side fell because people started doubting the payback period on hardware spend. The short side rose because people noticed that the most aggressive AI-disruption narrative had not shown up in anyone's numbers. A long-only fund loses money in that tape but survives it. A long/short book getting hit on both sides at 4x leverage does not. A margin call doesn't look at your annualized return. It looks at today's NAV. ## III. Six Days The timeline does most of the work: | Date | What happened | |---|---| | Early July | Fund peaks at $45 billion | | July 24 | Letter to investors calling the moment a "strategic entry window," inviting additional capital on August 1 | | July 30 | Microsoft earnings reprice AI capex; the fund is force-liquidated | | Within 36 hours of July 30 | Citadel takes the entire public-market book, about $16 billion, at a discount | | July 31 | SanDisk +25.99%, CoreWeave +21.51%, Bloom Energy +26.49% | Down 67% in the month of July. Assets held by the fund went from $45 billion in early July to about $10 billion on July 30. **He was liquidated at the low, and the low bounced more than twenty percent the next day.** That isn't bad luck. The mechanics of a margin call guarantee it happens near the low — the prime broker doesn't ask for money when you're comfortable. 4x leverage means your call has to be right every single day, not eventually. Aschenbrenner put part of the blame on short sellers targeting the fund's positions, saying they amplified the losses, and compared what happened to a "bank run." The analogy has something to it: a run doesn't require the bank to be insolvent, only that everyone wants their money at the same time. But a run can only happen if you borrowed short to do something long. He told investors the fund has now removed all leverage from the book. One detail deserves a separate look: who was on the other side of the trade. The buyer of that roughly $16 billion book was Ken Griffin's Citadel, at a discount. The next day the same stocks rose more than twenty percent. Which means the same assets, on the same fundamentals, were worth one price on July 30 and another price on July 31, and the difference wasn't in the companies. It was in who was getting margin-called. **That is the entire difference between a forced sale and a chosen purchase.** One party has to trade today. The other gets to trade today. The first one's price is set by the second. The 2000% cumulative return and the 439% first half look different in hindsight, too — not just as a record. Returns like that pull capital in as fast as capital can move, and the moment capital comes in hardest is usually the moment closest to the top. The $45 billion peak landed in early July, three weeks before the liquidation. **Size itself became part of the risk: the bigger the position, the harder it is to cut without moving the tape against yourself.** ## IV. His Call May Well Have Been Right That's the part of this that's hard to sit with. Look at the evidence on the demand side right now: - Amazon's Q2: AWS revenue of $42.2 billion, up 37% year over year, the fastest in 18 quarters; operating income $16.6 billion, up 64%; backlog of $496 billion. Jassy said 2026 capacity still won't be enough to meet all demand and 2027 probably won't either, and the company raised its 2026 capex budget from $200 billion to $220 billion. - Microsoft's Amy Hood said demand continues to exceed available supply. - Cook called memory chip pricing a "hundred year flood." **Not one of those three contradicts Aschenbrenner's essay.** Compute is tight, memory is tight, power is tight. The chain he drew in 2024 still holds on the demand side today. What would actually falsify the essay is demand collapsing on its own — large customers walking away from signed contracts, compute orders getting canceled, racks sitting empty. None of that has shown up. What got repriced in the week of July 30 wasn't demand. It was how much you have to spend to catch that demand. He didn't lose on direction. He lost on time. A correct long-term call, plus a short-term deadline you don't control, equals a wrong position. Leverage is that deadline — it compresses "will be right eventually" into "has to be right every day." On July 30 he didn't need to be wrong. He only needed to be late. ## V. This Is the Same Structure a Product Manager Lives In What I build is about as small as it gets, and none of it is comparable to $45 billion. But the structure is familiar, and it should be familiar to anyone who builds product. **A correct roadmap, plus a launch date that can't slip, plus a resource commitment that can't change — that is leverage.** You get a direction right, so you commit early: you hire, you kill other projects, you promise your boss a date. If that direction pays off six months later than you expected, your call is still right, but you're no longer in the room. The team gets reassigned, the project gets cut, and the person who eventually proves you correct isn't you. Two things here are worth pulling apart. **First, don't confuse "I called it right" with "I can survive until then."** The first runs on insight. The second runs on cash flow, on managing duration, on how much slack you left yourself. Most people spend 90% of their energy on the first and 10% on the second. **Second, be careful when your position and your narrative share a name.** Aschenbrenner naming the fund Situational Awareness was the strongest possible signal during fundraising — I believe this call so completely that I named myself after it. The same fact becomes the heaviest possible shackle in a drawdown: **cutting the position means admitting the essay was wrong.** The "strategic entry window" in the July 24 letter reads less like a man managing his risk than like a man protecting his thesis. Product managers have an exact equivalent. When a proposal becomes "your proposal," when a bet is written into your annual goals, when everyone on the team knows whose push this is — **the cost of adjusting it stops being a technical cost and becomes an identity cost.** From that moment on you start finding reasons for it instead of finding an exit from it. My own approach is crude: I write the call and the position in separate places. The call goes in one document, written as hard and as specific as I can make it, including what evidence would overturn it. The commitment goes in another, including how much I plan to pull back if it hasn't paid off by a certain date. They live apart so the second one isn't colored by the first one's feelings. It's not a sophisticated trick, but at least when I change my mind, I don't have to start by admitting I'm an idiot. ## VI. It Isn't Over The fund didn't blow up. What was force-liquidated was the public-market book. The private positions weren't sold, and those include equity in Anthropic. The leverage is fully removed. He still has money, still has positions, still has the essay. So the real open question is this: if 2027 actually arrives, if that chain of reasoning from the scaling laws is eventually paid off, the essay published on June 4, 2024 will be proven right — **but the investors whose positions were sold at a discount on July 30 will not be there to share in it.** I don't know how to land a conclusion on this one. A person can be right on direction and wrong on timing, and the market only settles the second one. That's not a new lesson. But watching it play out over six days at a scale of $45 billion is still hard to look away from. As for today — August 1, the day he had invited his investors to add capital. --- # Same Day: Microsoft Added $450 Billion in Market Cap, Meta's Free Cash Flow Fell to $784 Million URL: https://doaipm.com/en/blog/microsoft-cut-capex-and-grew-faster/ Published: 2026-07-31 Tags: Microsoft, Meta, Azure, Capex, Anthropic, OpenAI, Product Management, Tech Commentary Microsoft closed July 30 up 15.63% at $451.58, adding roughly $450 billion in market value in one day. That is the largest single-day market cap gain in US stock market history; the previous record was Nvidia's $441 billion on April 9, 2025. For Microsoft itself, it was the biggest one-day move since October 2008. Volume was close to 100 million shares, more than twice the daily average. The same day, Meta fell 9%, touched down 10.4% intraday, and lost about $130 billion in market value. Both reported after the close the night before. Both grew revenue. Meta's second-quarter revenue was $60.8 billion, up 28%. Microsoft's fiscal 2026 fourth-quarter revenue was $90 billion, up 18%. The one that fell grew faster. ## 1. The Market Wasn't Buying Those Three Points of Azure Growth Azure grew 43% this quarter, against 40% the quarter before. CFO Amy Hood guided next quarter to 45% — still accelerating. Azure crossed $100 billion in annualized revenue for the first time in fiscal 2026. Paid Copilot seats doubled, past 30 million. Net income was $35.8 billion, up 31%. Diluted EPS was $4.81, up 32%. For the full fiscal year, revenue was $331.8 billion and net income $133.7 billion. Good numbers, all of them, and nowhere near enough to explain $450 billion. Microsoft was worth close to $2.9 trillion before the move. A slightly-better-than-expected quarter does not buy a 15% single-day gap up. The actual thing was on another page of the same materials: **calendar 2026 capex guidance came down from the roughly $190 billion given in April to about $175 billion.** When Microsoft put out that $190 billion figure in April, the explanation attached to it was that memory prices were spiking. Three months later, the company says it doesn't need to spend that much. Azure accelerating and capex down $15 billion, in the same report. That combination hasn't appeared once in two years of AI narrative. The rule for the past two years was: you announce more spending, the stock goes up, because spending gets read as proof you captured demand. Whoever had the bigger capex number was assumed to have taken the future. Microsoft ran it backwards — it announced less spending and set a record. More important, Amy Hood said something else on the call: demand still exceeds available supply. **Spending less is not because they can't sell it.** Put those two sentences together and the meaning changes: it isn't that demand collapsed so they saved money. It's that the same demand can now be absorbed with less money. There's one more number on the demand side worth pulling out on its own: paid Copilot seats doubled, past 30 million. That number matters because it's seats — not monthly actives, not API calls. A seat means somebody paid per head, and most likely on an annual enterprise contract. The prettiest AI numbers of the past two years have been usage numbers: conversations, tokens, daily actives. They grow fast, and they sit several layers away from revenue. Paid seats don't have those layers. 30 million seats next to $175 billion of guided capex is what gives "spending less" somewhere to land. If Copilot seats were flat, the same sentence — we're cutting capex — would have read as retreat. ## 2. Where the $15 Billion Came From Amy Hood gave two reasons on the call. One is efficiency gains across the CPU and GPU fleets, getting more out of existing infrastructure. The other is process improvements in bringing new capacity online, shortening the cycle from decision-to-build to ready-to-use. That's the kind of answer product managers hate hearing. It isn't "we shipped a new feature." It's "we're getting more out of the hardware we already bought." That April figure of $190 billion says the same thing from the other side. The explanation at the time was spiking memory prices — meaning a meaningful share of that spend wasn't something Microsoft chose, it was pushed up from upstream. On a cost line you don't control, a company has two moves: accept the price increase and raise the budget, or squeeze more compute out of every dollar. In April it took the first. In July it delivered the second. This kind of work is the hardest thing to get resourced anywhere. There's no demo, no keynote, and it looks bad in an OKR — "improve utilization of existing clusters" reviewed next to "ship the AI assistant" loses almost every time. That day, the market priced it: **a $15 billion cut in spending, against roughly $450 billion in added market value.** I build a few small tools myself, at a scale with no relationship to any of these numbers, but the structure is the same. Whether a feature works after it ships and how much each call costs to run are two separate questions. Most people stop looking once the first one is answered. If the second one goes unanswered, the more successful the product gets, the faster the bill grows. What Microsoft signaled this quarter is that those two are no longer sequenced as "get it right first, optimize later." **In a category where unit costs are high enough to drag down the income statement, the cost structure is part of the product.** Those three points from 40% to 43% — if they'd been bought with $15 billion more hardware, the market would probably have read it as expensive life support. Getting them while spending $15 billion less is a completely different thing. ## 3. Part of the Savings Was Done by the Accountants That isn't the whole story. The same materials contained another move: starting in fiscal 2027, Microsoft is extending the depreciation life of data centers and office buildings from 15 years to 25 years, and reclassifying new data center leases from finance leases to operating leases. **That's an accounting change, not an efficiency gain.** The same servers, the same building, depreciated over 25 years instead, means a smaller charge hitting cost each year and a better-looking income statement. Moving leases from finance to operating also changes how they show up on the balance sheet and the cash flow statement. The roughly $175 billion figure is the number *after* this change in treatment. Blending this with the previous section turns the whole thing into cheerleading. They have to stay separate: one is running the machines fuller, the other is spreading the cost thinner. The market applauded both the same day, but only the first one is product capability. And the second leaves a real question behind: **can a data center actually last 25 years?** GPUs turn over roughly every three years. Cards bought a few years ago are no longer front-line for many training workloads. The 25-year assumption holds only if the buildings, the power and cooling, and the network backbone really do last that long, and the fast-depreciating portion is a small share of total assets. If the assumption doesn't hold, the charges were deferred, not erased. By fiscal 2028 and 2029, what has to be written down still gets written down. I don't know whether the assumption holds. But a company announcing it spent $15 billion less on the same day it extended asset depreciation by 10 years — those are two things worth remembering separately. ## 4. At Meta, the Money Went Out and the Cash Flow Went With It Meta's results weren't ugly. Second-quarter revenue was $60.8 billion, up 28% — well above Microsoft's 18%. The problem was in the lines below: - Diluted EPS of $6.18, 13.8% below consensus, mostly legal and severance costs - Capex of $31.1 billion for the quarter - **Free cash flow of $784 million** That $784 million is what actually broke the stock. A company doing $60.8 billion of revenue a quarter, with $784 million of free cash flow left, means capex ate nearly everything operations earned. Then management **raised** full-year 2026 capex guidance to $130–145 billion. That's the second consecutive raise. At first-quarter results, Meta lifted its 2026 AI spending expectation to $125–145 billion, and the stock fell that day too. This time the floor went from $125 billion to $130 billion. Two quarters, the same move, the same reaction. Three months in between, and the market was not persuaded. Same week, two companies gave opposite answers to the same question. Microsoft said it would spend $15 billion less and added $450 billion. Meta said it would spend more and lost $130 billion. Six months ago, Meta's move would most likely have been read as good news — the read then would have been "Zuckerberg is betting big." This time the read was "the money went in, nothing has come back yet." **What changed isn't Meta. It's the scoring rubric.** ## 5. Microsoft Owns a Piece of OpenAI and a Piece of Anthropic There's another set of numbers in this quarter's results that product managers should look at before the capex line. Microsoft holds about 27% of OpenAI. In November 2025, it also put $5 billion into Anthropic, and as part of the deal, Anthropic committed to purchase $30 billion of Azure services. This quarter: | Investment | This quarter | Full fiscal 2026 | |---|---|---| | Anthropic | **$3.2 billion gain** | — | | OpenAI | ~$600 million write-down (dragging EPS by ~7 cents) | $5 billion gain (contributing 67 cents of EPS) | The side it bet on for seven years and built up to 27% got written down this quarter. The side it only added at the end of 2025 contributed $3.2 billion. Look at one quarter and it's easy to write this up as "the backup saved the lead." Look at the full year and the OpenAI stake contributed $5 billion of gains and 67 cents of EPS in fiscal 2026 — still the larger number by an order of magnitude. A single quarter's write-down is volatility, not a verdict. The thing worth looking at isn't which bet paid better. It's the structure itself. $5 billion bought back a $30 billion Azure purchase commitment. The other side of that "investment" is customer acquisition — Microsoft bought not just a stake in Anthropic but a cloud order of known size. **Financially it's an investment. In product terms it's channel lock-in.** One level up: Microsoft is not betting on which model wins. It's a shareholder in both of the leading contenders and the cloud provider to both. OpenAI wins, Microsoft keeps the equity and the compute business. Anthropic wins, Microsoft keeps the equity and the compute business. Whichever breaks out, the traffic runs through Azure. What I build is as small as it gets, but the logic transfers. When the technical direction hasn't converged, picking a side is the most expensive move available — get it right and the upside is capped; get it wrong and everything you built on top of it is void. The steadier position is to be a required step on the path: **stay out of the question of who wins, and make sure that whoever wins has to come through you.** The cost is giving up the possibility of being the winner yourself. Microsoft is not going to be the company that builds the strongest model, and it has already conceded that by its actions. It picked a different position. ## 6. "How Much We're Spending on AI" Is No Longer a Selling Point Line up the day's events: - Microsoft cut capex, said demand still exceeds supply, and set a record on the stock - Meta raised capex, grew revenue faster, and fell 9% - Microsoft's stated reason for spending less was fleet utilization and delivery process improvements, not new features - The same day, it extended data center depreciation from 15 years to 25 **For two years, "how much we're spending on AI" was a marketing line. After that day, it became a number that requires an explanation.** For people building products, that shift isn't bad news. It means the work that gets resourced next won't only be the new features with a story attached. It also includes the things with no demo: getting inference costs down, getting cache hit rates up, scheduling idle compute out to something useful, changing how often a feature calls the model in the first place. That work used to be nearly impossible to prioritize. Its value just got publicly priced for the first time. As for the 25-year depreciation question, the answer arrives in fiscal 2028 and 2029. If those buildings and power systems really do last that long, today's math was honest. If they don't, the deferred portion still comes due in some year. What I'd rather know is next quarter. The 45% Azure guide was given on the premise that demand exceeds supply — and supply is exactly what's being held up by spending $15 billion less. How long both of those hold at once is the real test of whether this repricing stands. --- # He Gave Away 2.8 Trillion Parameters, Then Left One Gate in the License — Yang Zhilin, and Why He Still Isn't on This List URL: https://doaipm.com/en/blog/yang-zhilin-gave-away-the-weights/ Published: 2026-07-30 Tags: Yang Zhilin, Moonshot AI, Kimi K3, Open-Weight Models, Product Management, Tech Commentary On July 27, Moonshot AI uploaded the complete weights of Kimi K3 to Hugging Face. **1.42 TiB, 96 shards. 2.8 trillion total parameters, 104 billion active parameters, a 1-million-token context window.** Of its 93 layers, 69 run Moonshot's own Kimi Delta Attention and the other 24 use Gated MLA; there are 896 routed experts, 16 picked per token, plus 2 shared experts. This is the largest open-weight model in the world right now. **You can carry the whole thing home.** K3 went live as a service on July 16. Eleven days later the weights were given away. In those eleven days it scored 93.5 on GPQA Diamond, 67.5 on DeepSWE, 91.2 on BrowseComp, and 94.5 on MCPMark-Verified. The parameters and the benchmark numbers have been passed around all day. Three other things sitting in the same materials have barely been mentioned, and all three are product decisions. ## 1. The real design isn't in the model, it's in the license Everyone is saying "Kimi K3 went open source." **It is not MIT.** Moonshot wrote its own `Kimi K3 License`. The body of it does read like MIT — use it, copy it, modify it, distribute it, sublicense it, sell it, deploy it, fine-tune it, build derivative models on it. Then come two exceptions. **One: selling it as a service means coming back to the table.** If you turn it into model-as-a-service (MaaS) — giving third parties substantial control over inputs, parameters, or training data — and your company **clears $20 million in revenue over any consecutive 12 months**, you have to sign a separate agreement with Moonshot. **Two: get big and you carry the name.** A commercial product with **more than 100 million monthly active users, or more than $20 million in monthly revenue**, has to display "Kimi K3" prominently in the interface. **The exemptions**: purely internal use isn't restricted, and going through Moonshot's own products or certified inference partners isn't restricted either. Read those three passages together and the shape of the license comes out: > **Small players use it freely. The day you're making real money off it, either come back and split it, or run my billboard.** In my years doing product work, the open source strategies I've seen come in two flavors: genuinely wide open (and then you watch the cloud vendors monetize it while you get nothing), or a crippled edition released for show (and then nobody wants to build on it). **This license is a third road, and the incision is cut precisely.** The gate sits at the point where you have already made money — before that line, friction is zero, and individual developers, startups, and researchers help themselves. Anyone who crosses it had the money to pay in the first place. **Look again at that second clause, the one requiring "Kimi K3" to be shown prominently.** It doesn't collect cash. It collects something else — it turns every large product running on K3 into a Moonshot billboard. Open source buys ecosystem, ecosystem buys brand, brand buys pricing power, and that chain has been written into the license as legal text. There's a move here you can copy directly: **when you're trading "free" for scale, work out in advance the exact moment it stops being free, and write that moment down as a line the other side can check for themselves.** A vague free tier always ends in a mess of arguing. **The sequence was designed too.** K3 shipped as a hosted service on July 16; the weights came 11 days later. In those 11 days the benchmarks got run, the word of mouth built, daily revenue went up sixfold — **by the time the weights went out, it wasn't a model waiting to be validated, it was a model that had been proven.** Run it the other way around: weights and service on the same day, everybody self-hosts immediately, and you get neither those 11 days of revenue nor that first batch of real usage data. Open source isn't the same as not wanting money. It's **collecting the money and the data first, then putting the code out**. ## 2. They just took first place, and used the release materials to say they're behind This one is rarer. In their own release materials, Moonshot **concedes that K3's overall performance still trails Claude Fable 5 and GPT-5.6 Sol, and notes that there is a gap in user experience**. Not only that. They also volunteered a second point: **different models were evaluated using different agent harnesses** — meaning the benchmark environments are not fully comparable and the conclusions carry uncertainty. They had just built the largest open-weight model on earth and topped several hard evals, with the whole world watching, and **in their own announcement they said two things that work against them**: the experience still isn't as good as the competition, and the way these scores were compared is itself soft. Tesla's Q2 call yesterday was the mirror image. Musk said the humanoid robot demos going viral online right now are "mostly teleoperated or run off a pre-arranged script," that a robot capable of genuinely performing general tasks on its own hasn't appeared yet, and that Optimus will be the first — while reporting compiled that same week showed that at the 2024 Warner Bros. event, the Optimus units pouring drinks were being operated by engineers in motion-capture suits and VR headsets, and that at the Palo Alto headquarters a robot that falls over still needs a hoist to get it back on its feet. **Same week. One says everyone else's is remote control and scripts. The other says our experience still isn't as good as the competition's.** Both plays have their logic. Musk's expectation management has bought Tesla real time paid for in real money. Moonshot's candor may simply be because **it faces developers, and a developer can run the truth for himself inside a day** — in front of that audience, bragging costs more than it returns. But for a product manager, what you take away is the same either way: **every sentence you say gets checked by users against today's product.** The only variable is how fast your users can check it. The faster they can, the better honesty pays. **And developers are the fastest-checking users there are.** ## 3. Three times cheaper does not mean you spend a third as much The price is the part that got passed around the most, and it's the easiest to read wrong. K3's international pricing is **$3 per million input tokens and $15 per million output tokens**. Against Claude Fable 5's **$10 input and $50 output**, that's more than three times cheaper on paper. In the China region it's ¥2 for cached input, ¥20 uncached, and ¥100 for output. Then there's the detail that got buried in the press coverage: > **Multiple testers report that K3 burns more tokens than Fable to finish the same task.** That one line puts a discount on the "three times cheaper" above it. **Unit price is not cost. What you pay is the total spend to get one thing done, which equals unit price times the number of tokens it takes to get it done.** This is a textbook product trap, and it isn't confined to AI: **you optimized a number the user can see, and the price of it is hidden in a number the user can't compute.** I've walked into the same hole myself. Building my publishing pipeline, I was optimizing "time to publish on a single platform," squeezing every step shorter, and the whole thing got slower — because the shortened steps tripped the platforms' anti-automation, and the retry count went up. **The local metric looked better and end-to-end got worse.** So if you're evaluating a switch to K3, don't read the price list on the website. **Run your own real workload through it and total up the cost of one completed job.** That number may still favor K3 — the discount is large enough to survive a lot — but it should be a number you computed, not a number they printed. The demand is real: **after K3 launched, Moonshot's daily revenue rose at least sixfold; June ARR hit $300 million, up from $200 million in April.** ## 4. Who he is **Yang Zhilin, born 1992 in Shantou, Guangdong, top science scorer in Shantou's college entrance exam.** Undergrad at Tsinghua, transferred into the Yao Class in his second year, graduated first in his year in 2015. Then a PhD at Carnegie Mellon under Ruslan Salakhutdinov, Apple's first director of AI, and Google's William W. Cohen. **He is first author on both Transformer-XL and XLNet.** XLNet beat BERT on 20 tasks at the time and set best results on 18 of them. He is among the most cited NLP researchers under 35 in China. One more thing that has nothing to do with the technology and that I think matters: **he fronted a rock band at Tsinghua called Splay, as lead singer and lyricist.** In March 2023 he founded Moonshot AI with Zhou Xinyu and Wu Yuxin. In February 2024 Alibaba led a $1 billion round; in August, Tencent and Gaorong put in another $300 million. The gate in the license, the candor in the announcement, the trade-offs in pricing — none of those three are researcher thinking. A researcher's default move is to max the score and publish the paper. **Setting a precise commercial gate inside a license, and volunteering that you're behind the competition at your most triumphant moment, are product and business judgments.** A first author on XLNet doing both of those things at once is a lot more complicated than the label "technical guy with strong papers." ## 5. So why isn't he on this list I maintain a list called [The 100 Product Managers Who Changed the World](/en/rankings/), scored across six dimensions — vision, insight, taste, business, scale, originality — weighted into an OVR. **Yang Zhilin isn't on it.** The rule works like this: **the list has only 99 entries, and slot 100 is deliberately left blank, reserved for the reader.** To put someone in, you have to take someone else out. **Bumping a person off the list to chase today's news would produce a list that tracks the news cycle instead of product history.** So this piece doesn't add him. Not adding him isn't the same as saying he doesn't deserve it, and the standard should be stated out loud. **The person already on the list in the same bracket is Liang Wenfeng, OVR 91** (vision 95 / insight 84 / taste 86 / business 88 / scale 92 / originality 95). DeepSeek was an event that **changed the cost structure of the entire industry** — it broke the consensus that frontier models have to be expensive, and pricing and open source strategy moved worldwide in response. Holding the same ruler up to Yang Zhilin, my read is: **vision and originality are already enough to make the list; scale and business are still short of breath.** - **Originality**: turning a 2.8-trillion-parameter model into open weights is something nobody has done. That license with a revenue threshold in it is also new, and it will very likely get copied. - **Vision**: from Kimi's early bet on long context to today's bet on "open weights plus a commercial gate," the direction has been consistent and clear. - **Scale**: $300 million ARR and a sixfold jump in daily revenue is ferocious growth — but these are **numbers from a few months**, not a volume that has crossed a cycle. The scale dimension for the people already on the list measures how many people were affected, and for how many years. - **Business**: that license gate is beautifully designed, **and it has not been tested yet**. How many companies past $20 million in revenue actually come back and sign is something we'll know a year from now. **Put plainly: the most valuable part of his score is happening, not delivered.** That's the difference between a ranking and the news. The news records what is happening; a ranking records what has settled. **Putting him in today would be me betting on next year's result, not stating a fact.** **But this list is alive.** Allen Zhang and Steve Jobs probably aren't moving, while the middle and the back of it change constantly — when Linus was added back in, Will Wright came out. **I plan to re-rank the entire list in the second half of the year**, and by then what the K3 open-weight ecosystem actually grew, how much that commercial gate actually collected, and whether Moonshot survived the White House will all be plain facts. **If those have all been delivered by then, putting him on the list isn't a bet, it's a record.** The White House part is still hanging. After the K3 weights went out, **White House CTO Michael Kratsios publicly accused Moonshot of distilling Anthropic's Fable 5**, saying a distinction has to be drawn between normal AI distillation and "large-scale, covert, industrial-scale distillation designed to steal US proprietary technology"; Treasury Secretary Scott Bessent said the government "has the ability to impose sanctions." **An open-weight model that can draw out a country's CTO and its treasury secretary at the same time already carries weight.** But how that line plays out directly determines what his company's scale dimension is worth next year. ## Closing **"Free" was never the decision. "From what moment does it stop being free" is the decision.** Most people doing open source, freemium, or a trial period only work out the first half of that sentence and leave the second half fuzzy — and then either they can't collect when it's time to collect, or they throw up a paywall at the wrong moment and scare the ecosystem off. Yang Zhilin nailed that line down inside the license: **$20 million in revenue, 100 million monthly actives.** Whoever crosses it knows they crossed it. No negotiation, no arguing. I didn't think this carefully when I built SoloMD. It's free right now, MIT licensed, 551 stars, 21 external PRs — the ecosystem exists, but if someone turns it into a business tomorrow, I don't have a single line I can point them to. (All scores and rankings in this piece were produced by Claude (AI); the method is explained on the rankings page. The assessment of Yang Zhilin is this article's own analysis and is not counted in the list.) --- # CXMT Jumped 466% on Day One. The Company Has No Product Manager URL: https://doaipm.com/en/blog/cxmt-the-product-is-staying-alive/ Published: 2026-07-28 Tags: Tech Commentary, Product Management, Semiconductors, CXMT On July 27, ChangXin Memory Technologies listed on the STAR Market under the ticker 688825. The issue price was 8.66 yuan. It opened at 49.5, touched a 535.45% gain intraday, and closed at 49 — up 465.82%. Market capitalisation passed 3.28 trillion yuan. At the 8.66 issue price, the raise came to 57.919 billion yuan, beating SMIC to become the largest IPO in the history of the STAR Market. One winning lot of 500 shares was worth a little over 20,000 yuan on the day. Alibaba, as a shareholder, was sitting on more than 156.4 billion in paper gains. Those numbers circulated all day. Here is the thing I want to talk about instead: **the product this company makes is the kind that needs a product manager least of any in the industry.** ## First, the numbers - **Listing**: July 27, 2026, STAR Market, ticker 688825. Issue price 8.66 yuan, P/E of 308.92. - **Day one**: opened 49.5, peaked at +535.45%, closed +465.82%, market cap above 3.28 trillion. - **Raise**: 29.5 billion planned; 57.919 billion actual at the issue price (pre-greenshoe), the largest ever on the STAR Market. - **Position**: 7.67% of the global DRAM market — first in China, fourth worldwide. - **Product line**: DDR4, DDR5, LPDDR4X, LPDDR5/5X all in volume production; the only Chinese manufacturer making DDR5 at scale. - **HBM**: roughly 265,000 wafers/month total capacity, of which about 5,000 goes to HBM, targeting 30,000 by the end of 2026. Then the financials, which are what this company actually looks like: - 2023 non-GAAP net profit attributable to shareholders: **−16.752 billion**. - 2024: **−7.870 billion**, on revenue of 24.178 billion. - As of December 31, 2025, **cumulative losses of 36.65 billion**. - 2025 revenue 61.799 billion (+127% year over year), net profit 1.875 billion — it only crossed into the black the year before listing. - 2026 Q1: revenue 50.8 billion (**+719%**), net profit **24.762 billion**. - 2026 H1 guidance: revenue 110–120 billion, net profit 50–57 billion. A company that lost more than 30 billion over three years made 24.7 billion in one quarter. ## Where the 308.92x P/E comes from The 308.92 multiple gets quoted constantly as proof the stock is absurdly expensive. It is really an arithmetic problem. Divide the 579.2 billion issue-price market cap by the 1.875 billion of 2025 net profit and you get 308.9. So the three-hundred multiple is computed against **the loose change from its first profitable year**. The same prospectus guides to 50–57 billion of profit in the first half of 2026. Annualise that and the multiple is single digits. Both numbers are correct, and they are two orders of magnitude apart. **Which tells you the market was never pricing earnings. It was pricing the cycle.** ## DRAM is an industry with no product managers Back to the product. DRAM is one of the most thoroughly standardised industrial goods on the planet. The spec is fixed by JEDEC, an industry standards body: DDR5 is DDR5 — pin definitions, timing parameters, voltage, package dimensions. Samsung's, Hynix's, Micron's and CXMT's parts must be interchangeable if built to spec. That is the entire point of it. You don't think about who made the memory when you buy it. Which means the people who make DRAM hold none of the things a product manager normally holds: **No authority over the spec.** A committee defines it; you implement it. Think a parameter is badly designed? Doesn't matter. The standard is the standard. **No room for user insight.** The customers are PC makers, phone makers, server makers, and what they want is written on a purchase specification, word for word. There is no "what the user actually wants is…" here. **No taste to exercise.** What a memory module looks like, how it feels — none of it moves units. It gets slotted into a motherboard and is never looked at again. **No feature roadmap.** DDR6 is next; when it arrives and what it looks like is also up to the committee. So what is left? Process node, yield, capacity, capital expenditure, and timing the cycle. In the piece I wrote about Jensen Huang, he bet correctly that the GPU was a computing platform — that was a product judgment; he decided what the thing should be. In the piece on Wang Xingxing, Unitree's real product is the price curve — also a product decision; he decided what the thing should cost. **CXMT has neither. It does not set the spec, and it does not set the price.** ## Zhu Yiming, twenty years, not one product decision Zhu Yiming is chairman of CXMT and the founder of GigaDevice. Go through his twenty years and the key decisions are these: **2005, founds GigaDevice to make NOR Flash.** NOR Flash is a scrap-corner category in memory — small market, slow growth, not worth the majors' attention. He chose it not because it was good but because it was the one crack he could survive in. **2016, founds CXMT and moves to DRAM.** From the scrap corner into the staple crop. And where GigaDevice was fabless — design only, no fabs — CXMT went IDM: build your own plants, run your own wafers, do your own packaging and test. The heaviest, most capital-hungry, highest-failure-rate model in the industry. **July 2018, resigns as GigaDevice general manager** to run CXMT full time as chairman and CEO, pledging to take no salary until the company was profitable. He honoured that for seven years. **Partners with the Hefei municipal government.** A heavy-asset project that won't pay back for a decade is hard for commercial capital to carry alone; a local industrial fund underwrote it. **September 2019, the first independently designed 8Gb DDR4 enters production** — mainland China's DRAM industry goes from zero to one. Line those five up and **not one of them is a product decision**. Category selection, organisational form, funding structure, a commitment of time, an engineering grind — every one is a decision about how to stay alive, and none is a decision about what the thing should be. So: does CXMT have a super product manager? No. That seat is empty in this industry. **What sits in it is cycle judgment and capacity decisions.** ## What he designed was not a product, it was a structure One thing does deserve to be pulled out, because it is the most design-like move in the twenty years. Zhu Yiming holds two companies with opposite models. GigaDevice is fabless — chip design only, no fabs, wafers farmed out to foundries. Light on assets: small outlay, quick to turn, healthy cash flow, but a low ceiling, because you are forever at the mercy of somebody else's capacity. It took GigaDevice to the top tier domestically in NOR Flash and MCUs. CXMT is IDM — design, manufacturing, packaging and test, all in house. The heaviest model in semiconductors: a DRAM line costs tens of billions to build, and once built you still have to climb the yield curve. Fail to climb it and it is pure loss. Payback is measured in decades. **In 2016, the same year GigaDevice went public, he used that newly proven light company as his credential to go after the heaviest project available.** Put side by side, the two cards form a complementary design-plus-manufacturing structure: low-risk cash flow and design capability on one side, high-risk capacity and self-sufficiency on the other. The first proved he could build a company. The second was what he actually wanted to build. This move has more product instinct in it than any of the others — it decides not what a chip should be, but **what organisational form to use to carry something that won't pay back for ten years.** But note: it still isn't a product decision. It's an organisational one. Sometimes the most important judgment a product person makes isn't about the product at all. ## The product he actually built is endurance If we must name a product Zhu Yiming built, I think it is this: **he made "the company is still alive" into the product.** In an industry where somebody else defines the spec and a cycle defines the price, a company gets to decide exactly two things: how much to invest and when, and how long it can go without making money. Look at the timeline. Founded 2016. First DDR4 in 2019. Non-GAAP loss of 16.7 billion in 2023, another 7.8 billion in 2024, cumulative losses stacked to 36.65 billion. Seven years without salary. Then 2025 turns profitable on AI-driven memory price increases, Q1 2026 earns 24.7 billion, and on July 27 it lists at 3.28 trillion. **The turning point was not something he got right. It was that he was still on the field when the cycle arrived.** DRAM contract prices rose 93% to 98% quarter over quarter in Q1 2026. That increase did not come from CXMT's product getting better. It came from AI vacuuming up the industry's capacity: HBM's share of wafer starts goes from 18% at the end of 2025 to 22% at the end of 2026, and possibly near 30% by the end of 2027. Capacity flows to the higher-margin HBM, supply of the ordinary DRAM that goes into PCs and phones gets squeezed, and prices go up. UBS expects the tightness to last at least into the first half of 2028. Put differently: of that 24.7 billion quarter, how much CXMT earned and how much the cycle handed it is hard to separate. But one thing is certain: **if the 16.7 billion loss in 2023 had killed it, today's 3.28 trillion would have nothing to do with it.** That is what I mean by endurance as a product. It has a clear spec (survive to the turn), a clear cost (36.65 billion), a clear delivery date (seven years), and a real chance of failure. More than one Chinese memory project died on this road over the same period. And endurance has to be paid for by somebody. A three-year hole of more than 30 billion is not filled out of revenue; it was held up by a particular funding structure — the Hefei municipal industrial fund underneath it, then round after round of state and industrial capital. Commercial money struggles to carry a project alone that won't pay back for a decade and might get sanctioned to death along the way. There is not much money willing to sign that kind of long cheque. So strictly speaking, "he could endure" is not entirely his own capability; it is the outcome of an arrangement. What he got right was **accepting the price of that arrangement**: seven years without pay, his entire personal return deferred to the end. Post-listing his net worth is estimated in the 90-billion range, with more than 6,000 employees holding stock alongside him. That number exists because he took nothing for the first seven years. Both sides make it complete: he didn't carry this through on individual heroism, but he was the one who tied himself to the ship. ## So is this a great product manager I set one entry criterion for my own list of the 100 product managers who changed the world: only people who **created or invented a great commercial product**, judged on product decisions, not on the person. By that standard, Zhu Yiming probably doesn't make the list. He didn't originate a category — DRAM isn't his invention, DDR5 isn't his definition. What he did was take a standardised good that somebody else had already defined, whose spec was already frozen and whose players had been fighting for thirty years, and build it in China from zero to fourth in the world. The weight of that isn't in the product, it's in the industry. What it changed isn't what a memory module should be, it's whether China can make one. That is a national engineering achievement, and it is not the same category of thing as Jobs deciding the iPhone would have no keyboard, or Allen Zhang deciding WeChat would have no read receipts. **Praising both as though they were the same does neither any justice.** So I would put it this way: Zhu Yiming is not a super product manager. He is an engineer who bet twenty years on something certain to be hard and uncertain to work. That description is more accurate than "godfather of Chinese memory," and it deserves more respect than that phrase does. ## And some cold water **Cycles run both ways.** Thirty years of memory-industry history is alternating booms and crashes. This one is driven by AI capital expenditure — stronger than previous cycles, and more dependent on a single variable. UBS says tightness lasts into the first half of 2028; that is a forecast, not a fact. The memory of losing 36.6 billion over three years is only eighteen months old. **Between 308.92x and a single-digit multiple sits one assumption**: that 50–57 billion of first-half profit can be sustained. That is prospectus guidance, not a realised number. **7.67% share is fourth place, not joint first.** The top three are ahead of CXMT on capacity, process node and HBM progress. CXMT's HBM capacity is currently about 5,000 wafers a month, under 2% of its own total — and HBM is the most profitable part of this cycle. The target is 30,000 by year end, which is still the position of a company catching up. **A 465.82% first day says nothing about the company.** It says the issue price was set low, the float was scarce, and sentiment was hot. None of those three has anything to do with CXMT's yields. ## Closing On July 27, 2026, ChangXin Memory listed, rose 465.82% on the first day, reached 3.28 trillion in market cap, and raised 57.919 billion — the largest raise in STAR Market history. The same company lost 16.752 billion (non-GAAP) in 2023 and 7.870 billion in 2024, carrying cumulative losses of 36.65 billion by the end of 2025; its chairman drew no salary for seven years; it earned 24.762 billion in Q1 2026, in a quarter when DRAM contract prices rose 93% to 98%; it holds 7.67% of the global market, in fourth place; its HBM capacity is under 2% of its own output. The product it makes has its spec written by an international committee and its price set by a global cycle. --- # The 100 Product Managers Who Changed the World · No. 6 | Elon Musk: He Tore Out a Production Line That Was Making Money, to Make Room for a Robot That Can't Stand Up on Its Own URL: https://doaipm.com/en/blog/musk-tore-down-the-model-s-line/ Published: 2026-07-28 Tags: Elon Musk, Tesla, Optimus, 100 PMs Who Changed the World, Product Management, Tech Commentary On July 22, Tesla reported its second quarter of 2026. Revenue $28.236 billion, up 26% year over year — an all-time high. Operating income $398 million, down 57%. Capital expenditure more than doubled, to $5.8 billion. Free cash flow went negative: minus $1.1 billion. A company posting record revenue while its profit gets halved and then halved again usually means the market is beating it up — price cuts, lost share, costs out of control. That's not what happened here. **The profit that vanished is profit he spent himself.** There's a more specific line in the report: production of "other models" fell 34% year over year. "Other models" means Model S and Model X. Production didn't fall that hard because the cars stopped selling. It fell because **those two lines inside the Fremont plant were torn out**, and the floor space they freed up now holds the first-generation Optimus production line. And the robots about to come off it, by the company's own account, are not for sale. They'll first be used for "training data collection and further functional development." In plain terms: **he tore out two lines that were making money and handed the space to a product that doesn't make money yet and still needs a person to help it stand up.** When I had [Claude score the 100 product managers who changed the world](/en/rankings/), Musk came in at No. 6, OVR 95, with these six dimensions: **Vision 99 · Insight 88 · Taste 84 · Business 95 · Scale 97 · Originality 99.** Two 99s, a 97, and then two that visibly fall off the shelf: insight 88 and taste 84. That gap is what I want to talk about — because it and the torn-out production line are two sides of the same thing. ## Vision 99 and Originality 99: his real work is jamming product thinking into the physical world Start with his two hardest scores — which happen to be the two he ran through in public again this week. On July 24 at 18:50 US Eastern (6:50 the next morning Beijing time), Starship flew its thirteenth test flight. Three things were new. **It deployed 20 next-generation Starlink satellites in orbit for the first time**, establishing radio and laser links with every one of them. **It relit a Raptor engine in space**, the prerequisite capability for orbital changes and for coming home. And that 122-metre ship splashed down upright in the Indian Ocean — its gentlest water landing to date, and **the first time it made it all the way into the water without blowing up**. On the same mission, the first-stage booster's landing burn went wrong on the way back and it went into the Gulf of Mexico. That's what he looks like doing his job: one flight, half of it a historic first, half of it in the sea. The two 99s on vision and originality aren't awarded for "he succeeded." They're awarded for **repeatedly taking something everyone had agreed was impossible and turning it into a product that can be mass-produced and reused**. Before him, reusing a rocket was a romantic notion aerospace engineers kept to themselves; after him it's what Falcon 9 does on a weekly basis. Before him, satellite internet was a graveyard of bankrupt companies; after him it's broadband that several million people around the world are actually using. Before him, an electric car was an extension of a golf cart; after him it's a question of survival for hundred-year-old German automakers. **Product management as a job spends the overwhelming majority of its time circling around inside software, because software is cheap, iterates fast, and lets you undo your mistakes. He took the same set of instincts and jammed them into rockets, cars, batteries and satellites — places where one mistake costs hundreds of millions of dollars and several years.** On this whole list, only he and Jobs score 99 on vision and originality at the same time. ## Business 95: he'll trade a profit he already has for a product he might not get Back to the line that got torn out. On the earnings call, Musk said 2026 is a year of enormous capital expenditure, above $25 billion; he said "we are scaling up advanced infrastructure manufacturing capacity on a massive scale, and we believe this will be the largest such build-out in history." In my years doing product work I've seen the opposite scene far too many times: a business is still making money this year, so nobody dares touch it; even when everyone privately knows it's dead in three years, this year's number gets finished first. **"Trade a certain profit now for an uncertain future product" is the hardest and rarest class of decision a product manager ever makes** — because the person making it usually doesn't carry the consequences three years out, only this quarter's review. What Musk did this quarter is the extreme version of that decision. Model S and Model X are Tesla's flagship badge, old products with stable margins. Optimus is a new thing that doesn't even have a supply chain yet. He picked the second one. **Business 95 is for that**: he isn't merely "willing to bet," he can push current profit underwater while keeping the capital markets writing him cheques — net income $1.114 billion, down only 5% year over year, and global battery-electric vehicle production up 10% to more than 450,000 units. That's the chassis holding up the burn. Without it, tearing out a production line isn't vision, it's suicide. So why not higher? Because the books on this playbook aren't closed. Free cash flow has already gone negative, and Optimus, in his own words, is "the hardest product Tesla has ever tried to scale into mass production," with the biggest obstacle being that **there is simply no supply chain sitting there waiting** — the machine involves more than 10,000 unique parts, nearly every one of which has to be redefined, and no existing line can be picked up and dropped in. The bigger the bet, the less a full score is earned before it pays out. ## Insight 88: his most expensive lesson was treating his own conviction as the user's constraint Now the first docked score. Insight measures whether someone can see clearly what is actually happening in the real world — including seeing clearly that he is wrong. On autonomy, Musk bet on **pure vision**: eight cameras only, no lidar, no radar, no dependence on HD maps. The logic is genuinely elegant — a person drives with two eyes, so a strong enough visual neural network ought to be able to drive too; more sensors means higher cost and messier redundancy. The problem is that this logic has to clear regulation as well as physics. As of this July the comparison looks like this (numbers from a July 23 roundup by Sina Tech): - **Waymo**: roughly 4,000 autonomous vehicles in operation across 11 US cities, legally L4; more than 20 million cumulative rides by June 2026, and over 25% market share in San Francisco. - **Tesla Robotaxi**: 30 to 50 test vehicles (there were more than 200 back in January), of which only about 20 run driverless in Austin, and legally it's still L2 — the safety operator is the legal driver. Worse, the rules themselves are moving in the opposite direction. UN R57 requires redundant multi-modal perception for anything above L3. New Jersey's S1677 mandates radar for L3 and lidar for L4. China's 2026 standard requires forward detection of no less than 130 metres at 120 km/h, and bans relying on a single type of sensor. Meanwhile Tesla's own AI5 chip, due in 2027, supports multi-sensor fusion. **That's where insight 88 comes from: his conviction that "pure vision will work" was strong enough that for a long time he didn't take seriously the reality that even if it works, nobody is going to approve you to drive on the road.** Technical judgment and regulatory judgment are two different things; he is extraordinarily strong on the first and took a real beating on the second — and the cost is that the moment sensors go back on, most of the training data accumulated under pure vision has to be redone. I want to be clear about this: it isn't "he doesn't understand the technology." Quite the opposite — he's the kind of person whose technical conviction is strong enough to get a rocket built. **His conviction is his asset and also his bill.** The same trait carried him through a dozen explosions on Starship and cost him two extra years on Robotaxi. ## Taste 84: he guards the gate on specs, not on experience The second docked score is the one most easily misread, so I want to go slower here. Taste, on this list, isn't defined as "good aesthetics." It's **whether a person is willing and able to stand in front of the end user with his own judgment and hold the line on what counts as good**. Jobs's taste was not being able to live with a corner radius two pixels off. Allen Zhang's taste was everything he chose not to build into WeChat across ten years. Musk's gate isn't there. What he guards is specs and physical limits — seconds to 60, range, cost per kilogram to orbit, whether an engine can restart in space. On those he grinds down to the last decimal. But at the level of "how does it feel the moment a user touches it," his products have been split for a long time: the Model 3 interior shoved every physical button into a single centre screen, a decision people still love and curse to this day; the Cybertruck's stainless folded-plane shape is a textbook case of "I think this is what the future looks like" rather than "users need this." And this week handed us a sharper example. On the Q2 call, Musk said the humanoid robot demos going viral online right now are **mostly teleoperated or run off a pre-arranged script**, and that a robot capable of genuinely performing general tasks on its own hasn't appeared yet — he said Optimus will be the first. He also stressed that Optimus doesn't rely on pre-written programs but learns how to do things by observing its environment. That same week, a roundup from Wall Street Insight painted a different picture: at the 2024 Warner Bros. event, the Optimus units pouring drinks and chatting with guests were being operated backstage by engineers in motion-capture suits and VR headsets; at the Palo Alto headquarters, a robot that falls over still needs engineers with a hoist to get it back on its feet. Ken Goldberg, the roboticist at UC Berkeley, points out that the hard part of a dexterous hand isn't only the structure of the hand but the control system, environmental perception, and compensating for uncertainty. What Optimus is practising right now is basic tasks: sorting Lego bricks, folding clothes. **"Everyone else's is teleoperation and scripts, ours is real" — that sentence, coming out of a company whose own robot still needs teleoperation to demo and a person to pick it up when it falls, is precisely where taste 84 lands.** Taste isn't only aesthetics. It also includes **being honest about what your product actually is today**. Jobs oversold too, but what he oversold was something already in your hand, where one touch told you the difference. What Musk oversold this time is a capability still at the lab stage, and the way he did it was to call the competition a performance first. I don't think he's lying — on his time scale, he probably genuinely believes Optimus will be what he described a year from now. But every sentence a product manager says gets checked by users against today's product, not against the version running in his head for next year. ## So what did he actually give product managers One more thing from this week belongs alongside the rest. On July 23, The Economist released a 90-minute interview with Musk, recorded on July 20, staged in the main hall of Giga Texas — not an office in Silicon Valley, the factory floor. He said three things in it: AI may exceed the sum of human intelligence in **about five years**; the frontier AI companies should **peer-review each other** before models are released publicly; and the most likely outcome is "immense abundance for everyone," but the probability of catastrophic failure "is not zero." Put that interview next to the torn-out production line and his quarter makes complete sense: **if you genuinely believe AI will exceed the sum of human intelligence within five years, then still doing the math on Model X's quarterly gross margin is an absurd way to spend your attention.** It isn't that he doesn't know what tearing out a line costs. He's running the numbers on a different time scale. That's the thing about him most worth learning and hardest to learn: **the quality of a product decision depends on the length of the time scale you're doing the math on.** Most people's time scale is a quarter, a promotion, a funding round — which is why most people ship things that aren't ugly and aren't important either. And the other thing about him that's just as clear: **stretching the time scale doesn't buy you the right to be dishonest about the present.** The pure-vision bill and the teleoperated-Optimus bill both come due in reality eventually. Insight 88 and taste 84 are those two entries in the ledger. A man who can turn a rocket into a product, and who will also treat his own conviction as the user's constraint — OVR 95, No. 6. I think that's placed about right. He isn't the best product manager on this list. He's **the one who pushed the boundary of what product work can even be the furthest out**. (All scores and rankings in this piece were produced by Claude (AI); the method is explained on the rankings page.) --- # The TIME Cover Shows a $650,000 Mecha. Wang Xingxing's Best Seller Costs Under $5,000 URL: https://doaipm.com/en/blog/wang-xingxing-price-is-the-product/ Published: 2026-07-26 Tags: Tech Commentary, Product Management, Robotics, Unitree On July 23, the cover of TIME had a nine-foot machine standing on it: Unitree's GD01, the world's first mass-produced manned mecha, with a cockpit in its torso that a human climbs into. It fills almost the entire page. Founder Wang Xingxing stands beside it, made small by comparison. The cover line reads "The big robot moment." According to Global Times, this is the first time in eight years that a Chinese entrepreneur has appeared on a TIME cover. The GD01 starts at 3.9 million yuan. TIME puts it at $650,000. That number got passed around for a day. It is also the least important number in Unitree's product line. Turn a few pages into the piece and there is another set of figures worth more of a product person's attention. ## First, the numbers The hard data from TIME's reporting: - **Robot dogs**: from $45,000 down to under $2,000 over six years. - **G1 humanoid**: from $16,000 to $13,500 in eighteen months. - **R1**: under $5,000. - **GD01**: $650,000 (from 3.9 million yuan), roughly nine feet tall, half a ton, titanium-alloy limbs and a carbon-fibre shell, switches between two legs and four. - **Shipments**: over 5,500 units in 2025, more than a quarter of the global market. - **Revenue**: $62 million in Q1, up 68% year over year. - **Profit**: $6 million adjusted net, half what it was a year earlier. - **IPO**: filed in March at a $6 billion valuation. - **Customer mix**: only 9% to industrial applications, **74% to universities and research institutions**. Wang's own history is in the piece too: a native of Ningbo, he watched Marc Raibert's MIT robotics experiments at ten, built his first bipedal robot as a university freshman for under $30, and founded Unitree in 2016 at twenty-six. ## The one he sells most of never made the cover Line those prices up: $45,000 → under $2,000. $16,000 → $13,500. Then the R1 at under $5,000. That line is the product line. Anyone who has built hardware knows how hard cutting price actually is. It is not finance taking a slice off the quote sheet. It is pulling the whole machine apart and grinding down one component at a time — motors, reducers, controllers, structural parts — building each one in-house until the cost is under control, and only then earning the right to talk about price. Unitree holds all of those core components itself, not because vertical integration sounds impressive, but because this price curve cannot be walked down any other way. Taking a robot dog from $45,000 to under $2,000 in six years is a factor of 22. A cut on that scale swaps out the customer: a $45,000 machine is scientific equipment with a six-month procurement cycle; a $2,000 machine is consumer electronics you order after seeing it in a feed. Same machine — but once the price crosses a certain line, it becomes a different category. In the piece I wrote about Jensen Huang, the line was that the man selling shovels won on everyone else's gold rush. Wang took a different route: **make the shovel cheap enough that the prospectors no longer need approval to buy one.** Which makes the GD01 on the cover and what Unitree is actually doing two separate stories. ## He chose this road back in 2013 The price curve is not a strategy the company grew into. It was a technical choice Wang made in graduate school. He entered Shanghai University in 2013 for a master's in mechanical engineering. The mainstream path for quadruped robots then was hydraulics. Boston Dynamics' BigDog and Atlas were hydraulically driven — powerful, dynamically excellent, and also expensive, complex, leak-prone and hard to maintain. He skipped that road and went pure electric, using cheap off-the-shelf industrial brushless outrunner motors for the joints. In 2015, working alone on both the hardware and the control algorithms, he finished XDog, took second prize at the Shanghai Robot Design Competition, and published the electric-drive approach. Boston Dynamics published its own electric-drive work in 2016. There is a layer here that gets skipped. His stated reason for dropping hydraulics was lower engineering complexity and cost, with performance that held up well enough. **Translated: as far back as 2013, his first criterion for choosing a technology was whether it could ever get cheap enough to mass-produce.** A graduate student building a thesis project normally optimizes for impressive specs that publish well. He set himself a cost constraint instead. A bipedal robot for under $30 as a freshman. Industrial motors replacing hydraulics in grad school. A 22-fold cut on robot dogs over six years of running a company. That is one judgment repeated three times at three budget scales. By the time he sat for an interview in 2024, the number in the headline was a 99,000-yuan humanoid. Today the G1 is $13,500. Product people tend to treat pricing as a finance task — build the thing, then price it. Unitree runs it backwards: **price is a design input, not a design output.** Decide what it has to sell for before anyone will buy it, then work back to which motors, which materials, which control scheme. That ordering is what decides what the company looks like a decade later. ## The GD01's job is to be seen How many 3.9-million-yuan manned mechas can you actually sell? Wang has not bragged about that figure. TIME notes its design was inspired by UFC fights and the mechs in Avatar, and that its fists can punch through a wall. It is a showpiece. That is not an insult. A hardware company building a flagship it does not expect to ship in volume is a very old product move. The output is not orders, it is attention — and in hardware, attention converts. It buys press, it buys valuation, it buys the number investors are willing to put on a filing, it buys leverage in supply-chain negotiations, and it buys one extra mention of "Unitree" in the budget meeting at a university or a factory that will actually place an order. The marginal value of one mecha is that it makes the $5,000 R1 easier to sell. On the cover of TIME, that move hits peak efficiency: the whole world ran the campaign, and it cost nothing. The first Chinese entrepreneur on the cover in eight years — the label itself underwrites the company. Just don't read the showpiece as the results. The one on the cover and the ones carrying $62 million of quarterly revenue are two different batches of machines. ## The hardest number is 74% Of all those figures, the one to pull out is not $650,000. It is 74%. Unitree sells 74% of its shipments to universities and research institutions. Only 9% goes into industrial settings. In product terms: **its core users are not yet the people who put it to work. They are the people who study it.** Those two customers are nothing alike. A research institution buys a robot as a platform to run experiments, publish papers and give demos; its bar for reliability, cost per working hour and continuous run time is an order of magnitude below a factory's. A factory buys a robot and does arithmetic: how many stations it replaces, how long until payback, who repairs it, what an hour of downtime costs. TIME supplies that gap in numbers too: the G1 carries a 5-kilogram payload for 10 to 15 minutes at a stretch, and the machines run at 30% to 50% of human efficiency. Ten to fifteen minutes. That is the least flattering and most honest number in the report. It explains why 74% of the units went to labs — not because Unitree would rather not sell to factories, but because that endurance and that efficiency cannot yet hold down factory work. In the piece on Liang Wenfeng I made the point that DeepSeek could afford to go "narrow and deep" because High-Flyer had stockpiled the compute underneath it. Unitree's price curve has a precondition of its own: 74% of revenue comes from customers who are not that demanding, and that buys time. Research money is R&D budget — tolerant of failure, relaxed about metrics. The day industrial customers flip to the majority is the day Unitree faces the real product exam. So the thing to watch over the next few quarters is not how fast revenue climbs. It is whether that 9% turns into 20%. ## He gave the "ChatGPT moment" two to ten years himself On generality, Wang's line in TIME is that generalization is "the biggest headache for the entire global scientific community," and that embodied AI is "two to 10 years away from a ChatGPT moment." Two to ten years. That range is wide enough to mean "unknown," but it is more conservative than much of what his peers say — especially set against the mecha he just put on a cover. A founder who has captured the largest available platform in the world gives, in the same article, a floor of ten years. Generalization is stuck on data. Language models took off because the internet had been piling up text for decades, ready to use. Robots need motion data from the three-dimensional world — how to twist a bottle cap you have never seen, how not to slip on an unfamiliar floor, how to still recognize the same cup after the lighting changes. There is no internet to crawl for that. It has to be run out of real machines, one at a time. Which is another way to read those 5,500 shipments and that 74% lab share: every machine sold is out collecting data for Unitree in an unfamiliar environment. Selling more is itself the supply channel for training data. Wang offered one more call: within ten years, every household could have a small robot doing chores and care work. He did not attach a date to that one. ## And some cold water This is not a puff piece. TIME wrote the unflattering parts in, and they are worth copying out. **Safety**. A backdoor was found in the Go1 robot dog that allowed remote access to location and camera feeds. Clips of Unitree machines injuring children during performances have circulated. Wang's answer on things going wrong is "That won't happen. These robots are engineered with rigorous hardware constraints" — an answer that works for an engineer and explains rather less about the footage that already exists. **Subsidies and politics**. China has put more than $26 billion into related investment funds since late 2024; Unitree is one of Hangzhou's "Six Little Dragons," benefiting from a local $14 billion sci-tech fund, and humanoids have been designated a "disruptive innovation" by the Ministry of Industry and Information Technology. U.S. Representative John Moolenaar has accused Unitree of taking "generous state subsidies" and threatening American companies, and the House Guard Act would ban Chinese robots deemed security threats. Not a technical problem — but the single largest variable in Unitree's overseas market. **The books**. Revenue up 68% in Q1, adjusted net profit halved to $6 million in the same breath. Cutting price is paid for out of margin; the curve is not free. And a $6 billion valuation has not been tested by public markets yet. **Competition**. Tesla's Optimus, Figure 03 and Agility are all on the same track. Musk's line — "We don't see any significant competitors outside of China" — usually gets read as a compliment to Unitree, but it also says the real competition is inside China, and Unitree is not the only Chinese company building humanoids. The price comparison is the more interesting one. Tesla's target for the third-generation Optimus is a unit cost under $20,000, contingent on production ramping to a million units a year; outside estimates currently put its unit manufacturing cost at $50,000 to $100,000. So the price Tesla needs scale to reach is a price Unitree's G1 already sells at — $13,500, with the R1 lower still. There is another side to that. Tesla's $20,000 assumes a million units a year; Unitree shipped just over 5,500 in all of 2025. Unitree wins on today's price, Tesla is betting on the cost structure that follows volume. Those two eventually collide — if Optimus really does reach a million units a year, the cost advantage Unitree holds today through in-house components meets an opponent amortizing cost across scale. **Two orders of magnitude separate 5,500 units from a million, and that gap is itself a kind of cost capability.** Morgan Stanley's projections — 13 million humanoids by 2035, a billion by 2050, a $5 trillion industry — are the part of a story like this to take least seriously. They measure the current mood, not production capacity. ## Closing The machine on the July 23 cover of TIME costs 3.9 million yuan, is piloted from inside, and can punch through a wall. The product the same company sells most of goes for under $5,000. Robot dogs are down 22-fold in six years. Q1 revenue was $62 million while profit halved. 74% of units went to universities and labs, 9% to factories. The flagship humanoid carries 5 kilograms for 10 to 15 minutes. The founder's own timeline for embodied AI's "ChatGPT moment" is two to ten years. That is what the first Chinese entrepreneur on a TIME cover in eight years actually does for a living. --- # Liang Wenfeng Is Now the World's Richest AI Founder — By Doing the Opposite of Every Big Tech Giant URL: https://doaipm.com/en/blog/deepseek-founder-richest/ Published: 2026-07-25 Tags: big-tech watch, AI strategy, DeepSeek, product managers On July 14, the Bloomberg Billionaires Index updated one number: DeepSeek founder Liang Wenfeng is now personally worth $36 billion, roughly 243 billion yuan. In a single year, he more than doubled — up from $16.7 billion. Here's what that number means: he has surpassed OpenAI co-founder Brockman and Anthropic co-founder Amodei to become the single richest person in the world building AI foundation models. On China's rich list, he ranks eighth. But the number isn't the most interesting part. What's worth talking about is how he got here — by doing almost the exact opposite of every big tech giant. ## First, let's get the numbers straight - Bloomberg, July 14: Liang Wenfeng is worth $36 billion, more than double the $16.7 billion of a year ago. - This makes him the richest founder within "pure AI foundation-model companies" — not counting the sprawling conglomerates that do AI on the side. He's ahead of both Amodei and Brockman. - In June this year, DeepSeek closed its first external funding round, about 51 billion yuan (~$7.4B), valuing the company at 400 billion yuan (~$50 billion). Liang put in 20 billion yuan (~$3B) of his own money to follow the round. - After that round, he still holds roughly 78% of the company. - The valuation has been climbing hard: this April DeepSeek was worth only $10 billion; by June it hit $50 billion — a 5x jump in two months. - The company is already preparing for an IPO, planning to list on the mainland, with a filing possible as early as within 2026. DeepSeek is an AI team that grew out of the quant fund High-Flyer in 2023. In three years, it went from a fund's side project to this. ## Counterintuitive move #1: he raised a round and still holds 78% Start with a detail the headline number tends to bury: **after raising a round, Liang Wenfeng still holds 78%.** In the startup world, that's almost unheard of. A company at this scale would normally have gone through seven or eight rounds, with the founder's stake diluted down to the low teens — often single digits. Liang didn't do his first external round until June 2026, and even then he put in 20 billion yuan of his own to follow it. Behind this is a path completely different from the mainstream: **don't rely on frenzied fundraising and cash-burning to spread out.** Compare it to the standard script for star AI startups these past two years: raise a few billion dollars right out of the gate, spend it grabbing chips, grabbing people, burning compute, and prop up the valuation round after round. DeepSeek did the reverse — raised late, raised little, built the company's value up first, then let capital in. From a product-building standpoint, holding 78% doesn't say "he's rich." It says **he didn't turn the company into a sprawling operation that needs constant transfusions to survive.** If a thing can only be kept alive by burning cash forever, the founder's stake can't be held. Being able to hold it usually means the thing itself can stand on its own. ## Counterintuitive move #2: narrow and deep, no all-in-one bundle What DeepSeek does can be summed up in three words: narrow and deep. It does one thing — models — and it pushes engineering efficiency to the extreme: fewer chips, less money, training models that match the top tier, then open-sourcing them. It didn't go build the big-tech AI bundle: assistant, cloud, chips, agent platform, office suite… it touches none of it and just grinds on the model. This is a mirror image of the two giants I wrote about a few days ago. Writing about Alibaba, I said it was "doing the math" — breaking AI into a token supply chain, building the whole chain. Writing about Tencent, I said it was "paying tuition" — Yuanbao, Hunyuan, and Xiaowei inside WeChat, three identities fighting each other. The giants' predicament, much of the time, is the internal friction that comes from being "big and all-encompassing": the front line is too long, forces are scattered, and every piece has to be fed. DeepSeek flipped this around: **pull the front line to its narrowest, compress all the forces onto one point, then do something at that point that no one else can.** A team of a few hundred outran the single-point efficiency of giants with over a hundred thousand people. Every product manager understands this logic, but few dare to actually do it: **focus isn't cutting a few features — it's daring to let go of the vast majority of opportunities and bet on just one.** DeepSeek bet on "make the model strong, cheap, and open-source," and nothing else. ## Counterintuitive move #3: how did the free thing become the most valuable? There's an even more counterintuitive layer: DeepSeek's flagship product is open-source — you can download it for free and run it yourself. And yet this "free" thing is what propped up a valuation that 5x'd in two months. Intuitively, open source means no money. But what DeepSeek traded open source for is global adoption, developer mindshare, and ecosystem position — and in the end, all of that turned into valuation and the confidence to IPO. When I wrote about Tencent open-sourcing Hunyuan Hy3 and Alibaba open-sourcing Qwen, I made a point: **open source isn't a value, it's a phase-specific lever.** DeepSeek used that lever to the hilt — it has no big-tech distribution, no big-tech cloud; the only reason it can sit at the table is the global influence it bought with open source. Build the influence first; the money comes later. ## What this means for people who build products Set aside the "richest man" hook, and the most galvanizing part of Liang Wenfeng's story for those of us who build products is this: **it proves that in the AI era, a focused, ruthless, efficient small team really can beat the giants stacking money and headcount.** The leverage has changed. In the past, to build a big business, you first needed big capital, a big team, a big operation. Now, a small squad that has thought clearly about "do one thing, and do it to the extreme" can move a valuation in the hundreds of billions. Resources are no longer the only ticket — judgment and efficiency are. This is really the extreme version of what doaipm has always been saying: **focus, efficiency, and doing one thing others can't — it doesn't have to be big.** What one person, or one small team, can move today is more than at any time before. ## Also, a splash of cold water That said, don't mythologize it. DeepSeek's ability to go "narrow and deep" came with a precondition. Behind Liang Wenfeng is the quant fund High-Flyer, which stockpiled a large supply of compute chips in its early years. It's because he had money and chips as a foundation that he dared to do only models and skip everything else. This isn't a rags-to-riches fairy tale — it's a person already holding ammunition who chose a more focused way to fight. And that $36 billion is a paper number. It rests on a $50 billion valuation, and a valuation is investors' expectation — the IPO hasn't happened, nothing has truly been cashed out. As I said in the piece on the 3.17 million, a number on paper and money actually in your pocket are two different things. The wind shifts, and a valuation can shrink faster than it grew. But directionally, DeepSeek really did set an example for "small and refined": in an era where everyone believes "compute is everything, scale is the moat," it took the narrowest path and proved another possibility exists. DeepSeek grew out of a quant fund's side project in 2023; by April 2026 it was valued at $10 billion, by June at $50 billion; Liang Wenfeng's net worth doubled in a year to $36 billion, past the co-founders of OpenAI and Anthropic. And to this day, its main product is still an open-source model you can download for free and run yourself. --- # AI Tore Down the 'I Can't' Wall — the One Left Is in Your Head URL: https://doaipm.com/en/blog/impossible-wall/ Published: 2026-07-25 Tags: AI era, product managers, solo builders, doaipm I spent a long time these past few days staring at two sets of numbers. Three years ago, both would have sat squarely in the "impossible" column. The first set: DeepSeek, a team of a few hundred people that grew out of a quant fund, built a model that goes toe-to-toe with OpenAI. Its founder, Liang Wenfeng, is worth $36 billion this year — more than the co-founders of OpenAI and Anthropic. A few hundred people, out-building a company of 100,000+. The second set: the 2026 World AI Conference (WAIC) opened, for the first time, a dedicated zone for one-person companies — 180 solo projects on display. One person can also be a company. Three years ago, if you'd said either of these things to anyone in the industry, the answer would almost certainly have been the same word: impossible. Now they're just sitting there, plain for anyone to see. ## The "I can't" wall is being flattened Let me be concrete. The "impossible" that used to block an ordinary person was, in the vast majority of cases, not "this thing can't be done" — it was "I don't know how to do this." I can't write code, so I can't build an app. I don't understand design, so I can't produce a decent interface. I can't edit video, so I can't start an account. I can't set up a server, so I can't launch anything. Every "I can't" is a wall, holding an idea inside your head, unable to get out. These past two years, AI has flattened these "I can'ts" one by one. You can't write code — say what you want to Claude Code, and it writes it. You can't design — describe the look you want, and it produces it. You can't configure an environment, connect a database, or write a regex — things that used to take months of dedicated study are now a single sentence away. How concrete does this get? Someone who has never written a single line of code, as long as they can clearly describe the product in their head, can have — in an afternoon — a clickable, usable prototype running in a browser. That same thing, three years ago, was a small team plus several months of work. The core of the doaipm way of building, stripped down, is one sentence: **you say it, and AI builds it.** This isn't a metaphor. What you say becomes something that actually runs and that other people can use. The wall of technical skill really has fallen. ## Building software three years ago vs. building software today Let me make "the technical barrier has fallen" concrete, so you actually believe it. Three years ago, an outsider who wanted to build even the simplest little tool faced this path: either spend months learning a programming language until you could write something usable, or scrape together some money to hire a contractor or a programmer friend, translate your needs into words they understood, go back and forth on revisions, and wait weeks or even months — all while a single "that's technically hard to do" could talk you out of it at any moment. The vast majority of people fell apart at step one: "Me, an outsider, go learn to program? Forget it." Today's path: open an AI coding tool, describe in plain language what you want, and it writes it out and runs it right there. Not happy with it? Say in plain language "change this part to that," and it changes it. In an afternoon, you're holding something clickable and usable. You wrote not a single line of code the whole way — you may not even understand the code it wrote — and none of that stops you from getting the product built. This isn't the future tense. DeepSeek, Qwen, Claude — these models are sitting right there, free or very cheap, for anyone to use. The real barrier has moved from "do you have this craft" to "can you clearly say what you want." And the latter is a capability that everyone who understands their own needs already has. ## On the other side of the wall, there's actually quite a pile Don't think "one person building a product with AI" is still some rarity, something that only happens in the news. It's becoming everyday. Open any developer community and shares like these are more and more common: a person with no technical background built an automation script with AI and killed off a two-hour daily chore; a designer built a complete SaaS single-handedly, launched it themselves, and charges for it themselves; a content creator built a tool site serving only their own niche little need. Three years ago, each of these would have required an engineer, a budget, and a long stretch of waiting. Those 180 one-person companies at WAIC are mostly the same kind of thing: one person fixated on one specific itch, using AI to fill in the "craft" part, turning it into something launchable that people actually use. Individually they're small, but together they prove one thing — the line "one person can't build a complete product" is now void. Barely any of these people are geniuses. What sets them apart from most people usually isn't knowing more — it's that at the "I can't" moment, they turned and opened the chat box and typed the first sentence. ## But there's one wall AI can't knock down Once one wall falls, you quickly discover that the wall actually stopping most people has just moved to a new spot. Here's a concrete case: the same idea, put in front of two people. The first person thinks, "This must require knowing how to code, right? I can't, so forget it." And so this idea, to this day, is still in his head. The second person opens AI and types, "Help me build a little tool that automatically organizes my bookmarks." Today, that thing is running online, and he uses it himself every day. These two people both can't write code. The only difference is one thing: whether they typed that sentence into the chat box — whether they dared to first make something ugly. **Once the technical barrier drops away, what's exposed is another wall, and this wall is in your head:** I'm not good enough, this is too hard, I'm not ready yet, what if what I make is terrible and people laugh. AI can write the code for you. It can't press "start" for you. ## Many "impossibles" are just an assumption no one dared to touch Look at DeepSeek again. The wall it stepped over had an assumption pressed underneath it: "To build a top-tier large model, you have to do it like the big companies — stack the most chips, burn the most money." Almost everyone took this assumption as a given, and so "a small team can't build a good model" naturally became "impossible." What DeepSeek did was refuse to accept that assumption. It went and grinded on engineering efficiency, using fewer chips and less money to build a model on par with the top tier. Once it actually did it, that "impossible" wall, looked at in hindsight, wasn't a capability problem at all — it was that no one was willing to touch that default premise. Many of the "impossibles" in your head have this same structure. Take one apart and, pressed underneath, there's often an assumption you've never questioned: "you must have A before you can build this," "someone like me can't do B." Pull that assumption out on its own and ask, "Really?" — and the wall often loosens on its own. Stepping over a wall, much of the time, isn't about how hard you push; it's about whether you dare to question, "Says who this is impossible?" ## What the wall in your head looks like, taken apart This wall isn't one you get over with a "be brave." It's made of a few very specific bricks. Let me lay them out one by one, and you decide for yourself which one is blocking you. **Brick one: "I don't understand tech, so this isn't for me."** Three years ago this held up; now it doesn't. Tech is no longer the barrier. What's actually scarce is whether you understand users and can see the real problem clearly. A person who understands the problem but not the tech, paired with AI, runs far faster than a person who understands the tech but not the problem. Not knowing how to code is, today, even an advantage — you won't get tied down by "how do I implement this," and you can keep your eyes fixed on "what does the user actually want." **Brick two: "I'll do it once I'm ready."** You'll never be ready. In the AI era, the cost of building a prototype approaches zero, and the cost of revising a version approaches zero too. Under this cost structure, building an ugly thing that runs first will always beat thinking about it in your head for three months. In the three months you spend thinking, someone else has revised to version ten and gotten real feedback. **Brick three: "This is so simple, surely someone's already built it."** Built doesn't equal built well, and it certainly doesn't equal built the exact way you want it. And plenty of things are exactly right when you build them for your own specific little scenario that others look down on. Whether the market is big is an investor's problem; your own itch is worth an afternoon. **Brick four: "What if what I make turns out terrible?"** The first version should be terrible. The point of building a prototype that runs is to get it moving first and receive real reactions — not to nail perfection in one go. The "terrible" you're worried about is really you comparing it against the perfect finished product in your head. But the one in your head will never turn real on its own. In the real world, a usable sixty out of a hundred beats an imagined perfect hundred: the sixty can get feedback and grow version by version; the hundred can only sit in place and shine, useful to no one. ## Why it had to be now The "I can't" wall has actually always been there — it's been there for decades. Why did it fall these past two years, of all times? Because AI is the first tool where "you speak plain human language, and it goes to work." Before it, the things that claimed to help you clear the technical barrier — search, tutorials, low-code platforms — none of them actually got around the "you have to understand a bit of tech first" gate. You find the answer via search, but you still have to understand it; you use low-code, but you still have to grasp its logic. AI is different. It catches your plain language and spits out something that runs. That whole stretch of "you have to learn it first" — it swallowed the entire thing. What this brings isn't the barrier being lowered a bit; it's the nature of the barrier changing: from "learn a craft" to "clearly state a thing." A dislocation like this comes around maybe once in decades. And the biggest dividend of a dislocation period has always gone to the ones who dare to act first. Once everyone has caught on and is using it, the dividend flattens out. The people acting now are standing at the starting point of this dislocation. They aren't smarter — they just swallowed that "I can't" a little earlier and swapped in "let me try." ## What the people who stepped over it actually did I'm not going to give you a pep talk. Let me just tell you the common threads I've observed in the people who actually "stepped over" — check yourself against them. **One, translate "impossible" into a concrete sentence.** "I want to build a tool that automatically organizes my WeChat bookmarks" — as long as you can say it clearly, AI can start working. If you can't say it clearly, it means you haven't thought it through yet. That's "haven't figured it out," not "impossible." Don't confuse the two. **Two, ask for the ugliest running version first, not the complete solution.** A half-finished thing you can click and see results from today beats a perfect plan on paper. The first thing you want has only one standard: it has to run. **Three, change one thing at a time, let AI go step by step.** Don't have it build the whole thing in one shot — that way, when something breaks, you won't know where the error is. One small change at a time, take a look, then the next step. These points sound plain. The barrier to actually doing them is just this: open the chat box and type the first sentence. ## Back to those two sets of numbers from the start DeepSeek didn't win because it had the most compute. It's that a small handful of people were convinced "we can build a cheaper, smarter model," and then actually went and did it. Behind each of those 180 one-person companies at WAIC is a thought — "one person... could actually work?" — that someone pressed start on. What stands between you and these things is no longer whether AI can do it — AI already can, and can do far more than you think. What stands in the way is whether, the moment you say "I can't, I'm not good enough," a single action follows right behind it: saying it to AI. That wall has always been in your head. AI knocked down the one next to it and cleared away "I can't"; but the one that reads "I'm not good enough," it can't knock down — only you can step over that one. And the first step over is so small it barely looks like a step: open that chat box, and tell it, exactly as it is, the "impossible" you've been holding in. The rest, it takes from there. --- # 3.17 Million in Bonus, But Not Even the Freedom to Show It Off URL: https://doaipm.com/en/blog/salary-leak-watermark/ Published: 2026-07-24 Tags: big-tech watch, workplace, working life, high-voltage line Today a year-end-bonus screenshot went viral: at Tencent's WXG (Tencent's WeChat Group — the WeChat line), a project team lead had a year-end incentive of about 3.17 million yuan, 820,000 in cash and 2.35 million in stock. In 2025 he twice earned Tencent's top performance rating, Outstanding — a rating only the top 10–20% ever get. Put that in front of any working person, and it's a ceiling-level track record. Days later, that record was wiped to zero: fired, blacklisted, never to be rehired. The performance was flawless. What sank him was that he posted this screenshot. I stared at this news for a long while. The sharpest part isn't the bystander's regret of "3.17 million, gone." It's something else, something far more universal: **the money you break your back to earn — you don't even fully own the freedom to "show it off."** ## First, let's get the facts straight Here's the replay: - A pay screenshot went viral online first, clearly showing this lead's annual incentive of about 3.17 million yuan. - Someone noticed the screenshot carried Tencent's internal watermark — the moment it's posted, they can trace exactly whose hands it leaked from. - Tencent's anti-fraud investigation unit received the tip, stepped in, and found that during his tenure he had also sent sensitive company information to external parties, leaking it to outside platforms. - He was ruled to have crossed the third of "Tencent's high-voltage lines": information-security violations and disclosure of confidential material — which explicitly states that leaking or prying into salaries counts too. - Outcome: fired, blacklisted, never to be rehired. I won't dwell on the watermark part; one line covers it: **at Big Tech, every sensitive thing you see very likely carries a mark that points to you and you alone.** This time it just got acted out for everyone to see. What's really worth discussing is what comes after. ## Why a single pay screenshot goes viral While we're here, let me note why this news swept the entire internet within a day. Because it lit up, in one shot, the two most tangled-up emotions of working people. One is envy: 3.17 million — why him? We're all wage workers, so why does one person's incentive for a single year match what an ordinary person makes over a decade-plus of not eating and not drinking? The other is fear: if even a top performer — someone pulling 3.17 million, someone rated Outstanding twice in a row — can be wiped to zero overnight over a screenshot, then how secure is the little we ordinary people are clutching? Envy and fear tangled together become that "watching the spectacle, chills down your spine" kind of feeling. A piece of news that can make people envious, vindicated, and afraid all at once is bound to go viral. What it hits isn't one person's gossip — it's the raw nerve all working people share. ## Let's do the math first: of the 3.17 million, how much did he actually pocket Don't rush to envy that number. The 3.17 million breaks down as 820,000 in cash plus 2.35 million in stock. Stock is the bulk of it — three quarters of the total. Big Tech stock isn't something you can cash out the moment it's granted; it usually vests slowly over several years, in portions. You have to serve out your time honestly and stay out of trouble to collect it in full, year by year. Once someone is fired and thrown onto the blacklist, the unvested portion is basically wiped to zero. So the "3.17 million" that went viral was never a lump of cash landing in hand from the start. It's more like a rope tethering you for several years: to collect it in full, you have to swallow your words and hold the line over those years. In the end, this lead was kicked out before the rope had even finished paying out. What actually landed in his pocket was probably just that 820,000 in cash, plus a small slice already vested — a long way short of 3.17 million. Buried here is something a lot of people never think through: **Big Tech loves to pay high salaries in stock, and it's not just generosity.** Stock that vests over years is itself a design — using money not yet in hand to buy your good behavior for several years: don't dare talk loose, don't dare jump ship, don't dare cross the line. The industry calls this golden handcuffs. The cuffs are gold, full purity — but at the end of the day, they're handcuffs first. When you put them on, all you see is the gold; only when you go to take them off do you realize they've been tethering you the whole time. ## What kind of person pulls 3.17 million and gets Outstanding twice One more note, so you don't picture him as some connected type collecting money lying down. To make project team lead at WXG — Tencent's most profitable line — and land in the whole company's top 10–20% two years running, that's earned by real work. This kind of person is usually the backbone of the team, a core pillar the company is willing to pay a fortune for, deliberately binding them with stock in the hope of keeping them for the long haul. Precisely because he's that kind of person, "wiped to zero by one screenshot" reads all the more like an alarm bell: even the backbone, even the top performer — the moment the company decides you crossed the red line, gets cut just the same, no hesitation. At Big Tech, nobody's contribution is "great enough" to be exempt from the high-voltage line. The more important you are, the more it proves this line matters more than you do. ## Why leaking salary is the reddest of the red lines The "high-voltage line" isn't an adjective; it's a hard-coded system of red lines, and its defining feature is a **single-veto**: it doesn't matter how big your contribution or how high your performance — cross it and it's firing plus blacklisting, with no such thing as "offsetting the fault with the credit." One of those lines makes a lot of people freeze the first time they hear it: **leaking or prying into salaries is, in itself, a red line.** Why is salary so sensitive? Because Big Tech's pay systems are more complex than outsiders imagine: new hires out-earning veterans, same role different pay, stock vesting on different schedules — two people sitting side by side might have packages that differ by a factor of two. The moment this stuff gets laid out for side-by-side comparison, the internal sense of fairness, the poaching risk, the negotiating leverage all go haywire. So "don't show it, don't ask" is a near-universal red line at every Big Tech company — it's not that this one company is stingy. And here's the irony: he probably never felt there was anything wrong with showing off the money he'd worked so hard to earn. But under this system, showing your salary already crosses the line, and the information the screenshot carried out crossed an even heavier one. A single forward lit two fuses at once. ## The tangle: the industry runs on "showing off," but an individual who shows off dies There's another contradiction here, one rarely laid out plainly. The entire hiring market runs precisely on "showing off." Maimai, Zhiyan, all kinds of offer-sharing — everyone shows their package, compares packages, benchmarks total comp — that's how you know what the same role is worth on the market and how much you should ask for. You could say the only sliver of bargaining power working people have on pay is pieced together, bit by bit, from exactly this information that "was never supposed to be shown." But inside the company, showing your salary is a hard kill line. So working people get stretched across the middle: **you can only learn your own worth from someone else secretly showing theirs; the moment you show yours, you might get zeroed out.** The information here is utterly asymmetric — the company knows every person's package cold, while you can only guess from the odd leaked screenshot. In a sense, this lead stepped right onto that fault line. The picture he showed off is a rare pricing reference for peers, and a leak the company must stamp out. Same picture, two fates. ## What high pay buys isn't just your time We default to one assumption: the company pays for my time and output, and once the money's in hand, it's mine, to do with as I please. This news is a reminder to everyone: **Big Tech's high pay also buys away your right to dispose of that gain, and your right to speak about it.** The 3.17 million landed, but you can't show it, can't discuss it, can't even feel a little proud about it on your feed. This money is written under your name, yet fenced in by an invisible red line. There's an even more gut-punch of a contrast: **the better your performance and the higher your rank, the more dangerous you are — not the safer.** Two Outstandings mean he saw more, had wider access, and held more valuable sensitive information; the moment something goes wrong, the blast radius is bigger too. So the more core the person, the tighter the information string gets watched. In this logic, top performance isn't a talisman — it's, to a degree, a larger risk exposure. We always imagine "grinding your way to high pay, grinding your way to high rank" as the road to freedom. Only when you actually reach that spot do you discover the ropes are more, and tighter, than before. ## Why it's Big Tech, specifically, that clamps down this hard Show your salary at a small company, and at most the boss is quietly unhappy — it rarely reaches "firing plus never-to-be-rehired." At Big Tech, why come down so hard? At bottom it's two words: worth money. The bigger the company, the more valuable the information in its hands — one piece of internal data can move the stock price, one playbook can be copied by competitors, one leak might sit atop the privacy of over a billion users. Once information can't be held, what's lost is the whole business. So it has to weld shut, as far as possible, the mouth of everyone who can touch sensitive information. This is "the price of scale": you enjoy Big Tech's platform, résumé, and high pay, and in exchange you have to accept Big-Tech-grade constraints. The two ends of the scale have been tied together since the day you signed that contract. ## This isn't a dilemma exclusive to top performers You might think 3.17 million has nothing to do with you. But the thing that boxed him in boxes in every working person — his scale just makes the cost especially jarring. Pull the lens back to ordinary people, and the same class of "information red line" is actually right beside us every day: - Sending an offer screenshot into a job-hunting group chat for advice, only to have the position snatched by someone. - Venting anonymously about the company on Maimai, then getting doxxed by a colleague who followed a few details. - Before leaving, casually packing up the decks and spreadsheets you made to use as a portfolio at your next job. - Explaining the last company's playbook and data in full during an interview, to prove your chops. - A line on your feed — "finally got the year-end bonus!" — with the amount attached. - Treating an unannounced project or org reshuffle as gossip to tell friends outside. - Saving internal group chats and Feishu (Lark) docs on your phone, screenshotting and forwarding on a whim. Almost everyone has done these, or at least toyed with the idea. Not one of them is "deliberate sabotage" — they're all convenience, laziness, a little showing off, or just wanting some advice. The red line gets crossed precisely in these "didn't think much of it" moments — you feel it's sharing, the system rules it a leak. The only difference: he crossed the line and lost 3.17 million plus his ticket into the entire Big Tech circle; you cross it and might just get a talking-to from HR. **The constraint is the same set — only the price tag differs.** While I'm at it, a word on the weight of "never to be rehired." Big Tech companies share, to some degree, their lists of dishonesty and fraud, and background checks at same-tier companies can turn it up easily. Once you're on it, the door to moving up basically shuts along with it. What he lost isn't one job — it's the entry ticket to the whole circle. The former can be earned back; the latter rarely comes again. ## So how should we actually read this Feeling sorry for him is only human. But if we really want a "lesson" that lands on us, it's plain: **at Big Tech, treat everything internal you see or receive as "real-name" by default** — including your own pay slip. Before you screenshot or forward, spend one more second asking whether it goes out with your name on it. More worth chewing on than that is the awkwardness: we treat high pay as the freedom of "finally making it out," and at the same time discover that the further up you climb, the less you actually get to decide for yourself. That 2.35 million in unvested stock is, at bottom, the price tag on this awkwardness. We're used to measuring a job's worth by "total comp" — the bigger the number, the more settled we feel. What this news tears open is the line of fine print behind that string of digits that nobody reads to you: whether you can collect this money in full, spend it with peace of mind, or mention it proudly in front of others — all of it hinges on not crossing the line once in the years ahead. **Total comp is shown to you; the constraints are there to bind you.** Both were handed to you together, from the day you signed the contract. 3.17 million in cash and stock, two Outstandings, years of accumulated professional credit — all wiped to zero by a single screenshot he thought was his own. --- # Alibaba Is Doing the Math, Tencent Is Paying Tuition: A PM's Read on the Big-Tech AI Split URL: https://doaipm.com/en/blog/tencent-alibaba-ai/ Published: 2026-07-23 Tags: product managers, AI strategy, big-tech watch, AI commercialization In one quarter, Alibaba's cloud AI revenue crossed 30% of external revenue for the first time. In that same quarter, Tencent's new AI products lost about 8.8 billion yuan — roughly 35 billion yuan annualized. Both companies chant "all in AI." Both poured in over 100 billion yuan. So how did the gap widen this far? In this piece I want to lay the two companies' first-half-2026 AI moves side by side, from a product manager's angle. Here's the conclusion up front: neither model is weak, and neither loses on benchmarks. What actually opened the gap is something far more mundane — **can you say, in one sentence, what this AI is for and how the money adds up?** Alibaba answered that question crisply from day one; Tencent took a long detour, paid a hefty tuition, and only recently found the door. ## First, two timelines Let me lay out both companies' first-half-2026 moves. No commentary — feel it for yourself. Alibaba's side: - On March 16, it merged core assets — Tongyi Lab, MaaS, and Qwen — into a new business group called "Token Hub" (Alibaba Token Hub, or ATH), with the CEO taking direct command. The pitch is one line: make tokens, supply tokens, use tokens. - It set a five-year goal: grow cloud and AI commercial revenue from 100 billion yuan this year to 100 billion USD, a compound growth rate of roughly 47%. - In the Q1 report, Alibaba Cloud's AI-related revenue crossed 30% of external revenue for the first time. Over the same period it launched its self-designed Zhenwu M890 AI chip, a 128-card server, the flagship model Qwen3.7-Max, and Qianwen Cloud for agents. Tencent's side: - In March, it dissolved the AI Lab it had run for ten years, folding the whole team into the Hunyuan group. In the same month, its self-built desktop office agent "WorkBuddy" launched. - The standalone app Yuanbao had about 109 million monthly active users, trailing Doubao (315 million) and Qwen (202 million). - WorkBuddy climbed fast, hitting No.1 in China for monthly visits among AI office agents on PC and overshadowing the standalone assistant Yuanbao. Media headlines put it plainly: "Yuanbao out of favor, WorkBuddy takes the baton." - In June, WeChat's own agent "Xiaowei" entered gray-release testing: swipe right from the main screen to enter, call mini programs in natural language, and go all the way through to placing an order. WeCom's "Dayuan" also entered testing. - On July 6, it open-sourced Hunyuan Hy3 under the Apache 2.0 license — 295 billion total params, 21 billion active params, 256K context. In the same window, one company is talking revenue share on its earnings call, and the other is explaining losses while quietly switching ships. This isn't a gap in model capability. The gap is somewhere else. Read on. ## Alibaba: it turned AI into a business you can reconcile The name "Token Hub" is itself the answer. ATH's logic is to break AI into a supply chain: make tokens at the bottom (train models), supply tokens in the middle (deliver via cloud and API), use tokens at the top (agents and applications). The whole company revolves around one word — "token." The upside is something any product manager sees at a glance: **it gave the whole company a North Star metric you can reconcile.** Whether token consumption grew, how far cloud revenue penetration has reached — you can put those on the table quarter by quarter. Usage-based billing; the books are clean. Qwen's open-sourcing is the entry point of this whole game. By April 2026, cumulative global downloads of the Qwen series approached 1 billion, over half of all open-source model downloads worldwide. Open source is free; the point is to get developers everywhere using it and getting used to it first — to lay down "supply" before anything else. The math here works like this: open-sourcing the model spreads adoption, and the inference demand that actually runs ends up landing on Alibaba's own cloud, turning into a real compute bill. A developer running free Qwen locally today will, when their business scales tomorrow, most likely buy the cloud's API and compute. **Open source is customer acquisition; cloud is the cash register.** Then comes the close. In the first half of this year, Alibaba's moves shifted clearly from "everything free" to "the best model costs money": the flagship Qwen3.7-Max went closed-source, with its API opened only on Alibaba's own cloud. Open source captures the market, closed source and cloud collect the rent — the two layers lock tightly together. For people who build products, there's a lesson here: **open source isn't a value, it's a phase-specific lever.** Alibaba traded open source for global developer adoption at scale, and once the ecosystem took off, it monetized through the closed-source flagship and cloud infrastructure. It knows exactly what it's trading for at every step. Worth a side note: Alipay, in the same camp as Alibaba, hasn't been idle either. The AI version of Alipay, "Abao," opened public beta in early July — ship a package, hail a ride, order food, check your housing fund, all in one sentence — and behind it is the same "call mini programs + wire up payment" playbook. In other words, Alibaba's camp isn't only in the picks-and-shovels infrastructure business; it's also bet on the consumer-facing super-app agent line. ## Tencent: it took the Yuanbao fall to figure out where its AI should actually grow Tencent's foundation is the thickest in all of China: WeChat's billion-plus users, the relationship graph, payments, mini programs — no rival comes close on distribution. In the AI era, this should have been a crushing advantage. Yet its most instructive lesson of the first half is precisely a pitfall: Yuanbao. Yuanbao is Tencent's card for building a "standalone general-purpose assistant" to go up against Doubao. But at about 109 million monthly actives, it's left far behind Doubao (315 million). Why can't Yuanbao catch up? Behind Doubao is Douyin, pouring users in on the massive traffic of short video and willing to eat losses to subsidize. Tencent has traffic too, but its traffic lives inside WeChat — and WeChat is exactly the place it least dares to mess with. As a standalone app, Yuanbao amounts to Tencent going head-to-head with Doubao on a battlefield where it holds no advantage (standalone assistants), with a slow start to boot — hard to catch up. Here's the interesting part: while Yuanbao fell behind, Tencent's AI actually turned things around in two other directions — and what they have in common is that neither tries to be a "standalone assistant." Instead, they **grow into specific scenarios.** One is WorkBuddy, the desktop office agent that launched in March. Its monthly visits quickly reached No.1 in China among AI-native office agents; per-user token consumption grew tenfold in three months, and retention held around 60%. The media put it bluntly: "Yuanbao out of favor, WorkBuddy takes the baton." The other is Xiaowei, inside WeChat. After hesitating for most of a year, in June 2026 WeChat's own agent "Xiaowei" entered gray-release testing. There's a signal here worth noting: the card being played isn't the standalone Yuanbao — it's WeChat itself stepping onto the field. Yuanbao, Hunyuan, and WeChat AI once fought each other as three competing identities; now the direction has converged — **give up the standalone assistant, and embed AI into the scenarios it's already strong in (WorkBuddy for the office, Xiaowei for WeChat).** The cost is that Yuanbao's year-plus of investment basically went down the drain — that's the tuition. The WeChat agent is the heaviest card in the deck: let AI place orders for you inside WeChat, call mini programs, handle service accounts. Done right, that's a closed loop of transactions, relationship graph, and payment — terrifyingly powerful. That it waited this long before moving also shows how cautious Tencent is about touching WeChat's home turf — this is the classic innovator's dilemma: **the more valuable the home turf, the less you dare to touch it with AI that isn't mature yet.** But this card isn't Tencent's alone. Alipay's "Abao," in the same camp as Alibaba, entered public beta in early July, holding the same payment-plus-mini-program hand — and on the super-app agent path, it moved a step ahead of even WeChat's Xiaowei. The very distribution edge that looks most like Tencent's moat is being contested by a rival of the same scale. ## One move made in reverse exposes the stage each company is at Here's the interesting thing: in the first half of 2026, the two companies moved in exactly opposite directions on open source. Alibaba's flagship is pulling toward closed source, with Qwen3.7-Max cloud-only. Tencent's Hunyuan is pushing toward open source, with Hy3 straight to Apache 2.0. The same decision, opposite directions, yet both right — because the two are at different stages: - Alibaba already has an ecosystem and adoption at scale; now it's shifting from "capturing" to "collecting rent," at the point of closing off and monetizing. - Tencent hasn't yet built developer mindshare; using open source to quickly trade for ecosystem position and developer goodwill is buying an entry ticket with open source. Put these two moves side by side, and they explain more than any model benchmark could: **to judge a company's AI strategy, don't just look at whether it open-sources — look at what it's actually trading for right now.** ## The six dimensions a product manager watches Lay the two on one table and the differences jump out. | Dimension | Alibaba | Tencent | |---|---|---| | North Star metric | Token consumption / cloud revenue penetration, clear and optimizable | Shifted from Yuanbao MAU to scenario-agent usage, just recalibrated | | Business model | Usage-based billing, books add up cleanly | Ads / office subscriptions / payment loop, path still being explored, currently still in the red | | Who it serves | Developers, enterprises (B2B PaaS logic) | Consumers, workplace, social (consumer scenario logic) | | Growth engine | Supply side: developer adoption drives tokens | Demand side: embed agents into the office and WeChat | | Moat | Cloud infrastructure + global open-source share | Relationship graph + payment + mini programs, but Alipay is contesting the same entry point | | Organization | ATH concentrates forces, CEO in command, metric pressure | Dissolved AI Lab into Hunyuan; Yuanbao exits, resources bet on scenario agents | To sum up the table in one line: **Alibaba is selling means of production; Tencent is stuffing AI into the scenarios of life and work.** Alibaba's books added up from the start; Tencent's ledger took over a year to turn to the right page. ## The next two or three years for both **Alibaba: the story is more "financeable," and it's most likely the steadiest runner on the "AI as infrastructure" line.** If tokens really do become the "new electricity," as many say, Alibaba is holding the grid operator's seat. Add Qwen's nearly 1 billion global downloads, and it has also banked a chip no one else has in going global and in worldwide developer mindshare. And the Alipay "Abao" consumer super-app agent line gives Alibaba's camp a second hand beyond infrastructure. Its risks sit in three places: first, a token price war — everyone is cutting prices, and margins will get ground thin; second, whether its self-designed chips can fill the gap left by restricted Nvidia access — the Zhenwu M890 is only just starting; third, the tension of betting on both open and closed source — the further the flagship pulls toward closed source, the more it has to eventually confront head-on whether the open-source community's pull gets weakened. But the direction is self-consistent and the metrics form a closed loop — that's its biggest certainty. **Tencent: stop chasing one hit general-purpose assistant as the goal; its opportunity is in "scenario agents."** WorkBuddy has already proven one thing: as long as Tencent's AI grows into a specific scenario (the office), it can build volume — and build all the way to No.1. The real decider next is **whether Xiaowei inside WeChat can move from gray-release testing to full rollout and truly embed into WeChat's main entry.** If it works, once WeChat's loop of transactions plus relationship graph plus payment starts spinning, the flywheel is fearsome; if it doesn't, the best distribution in all of China will spin its wheels in internal friction. And there's now an added variable: Alipay's "Abao" has jumped the gun on the super-app agent, so this fight is no longer WeChat racing against itself. ## What people who build products can take from this comparison Set the two companies aside and come back to building our own products. The biggest contrast between them is really two sides of the same basic homework. Alibaba runs steady not because its model is the strongest, but because from the start it translated AI into a metric the whole company could reconcile — so everyone knew where to push. Tencent took a detour not because its tech is weak, but because it spent over a year, and burned a Yuanbao, before it figured out "what my AI product actually is" — not building a standalone general-purpose assistant, but embedding into the scenarios it was already strong in. Positioning is the kind of thing where, if you choose wrong, all the resources in the world are just tuition; being able to admit the mistake and pull resources back to the right direction is itself a skill. Yuanbao burned over a year and is still chasing MAU; WorkBuddy launched and hit No.1 among office agents in three months; Xiaowei bet the chips back on WeChat itself. Three AI products from the same company — one still paying tuition, two turning things around inside a scenario. --- # This Round of Big Tech Layoffs Is Hunting the Product Managers Who Just Pass Messages Along URL: https://doaipm.com/en/blog/layoffs-cut-the-messenger/ Published: 2026-07-22 Tags: AI layoffs, Big Tech layoffs, product managers, AI-era PM, tech commentary Line up a few things that actually happened this month. According to Layoffs.fyi, as of July 20, the tech industry has logged 302 layoff events in 2026, affecting roughly 200,000 people — about a thousand jobs lost every single day. More than half of those layoff notices name-checked AI in one way or another. Microsoft cut 4,800 people. Sales, consulting, Xbox — nobody got a pass. But Microsoft's HR chief, Amy Coleman, said something that ran exactly counter to the crowd: **"The roles eliminated today were not replaced by AI."** Meta is even more tied in knots: it cut around 8,000 people — a tenth of its workforce — and at the same time moved about 7,000 people into newly created AI-priority roles. One hand pushing people out the door, the other stuffing people into the AI division. Put these side by side and something strange jumps out: **everybody is saying "AI took the jobs," but when you drill into any one company, even they can't tell you what AI actually replaced.** Microsoft is busy distancing itself; Meta is cutting and hiring in the same breath. Only the bystanders are dead certain the machines stole the paychecks. So what actually got cut? ## What's being cut isn't headcount — it's a layer Lay out the roles that got mass-cut over the past six months and you'll notice an unflashy thing they have in common. The deepest cuts rarely land on the best doers on the front line, or on the top specialists. They land on the middle layer — the people who translate the boss's intent into tasks, roll up the front line's progress into reports, and align information between two departments. When GitLab restructured for "the agentic AI era," it said plainly that what it cut was a couple of layers of management. What Meta has done over and over these past two years goes by the name "flattening." The core work of these roles boils down to one thing: **passing messages.** Take what A said, translate it into something B can execute; take what B did, roll it up into progress A can understand. Meetings, alignment, chat threads, weekly reports, chasing schedules — at the end of a day, the part that actually creates something from nothing is tiny. Most of the time is spent shuttling information between people. And shuttling information happens to be exactly the work AI does most smoothly and takes over first. A model that never tires, can read every requirement doc at once, sync every project's status, and generate a structured report on demand lands precisely on the message-passing layer. So what this round of layoffs is really erasing isn't a bunch of individual heads — it's an entire layer of "relay" function. > Once passing messages can be handed to a tireless model, the roles that live on passing messages become the line item on the layoff list that needs the least explaining — even Microsoft can't be bothered to chalk it up as an AI win. ## Product managers are one of the most message-dense roles there is Right about here, product managers should be shifting in their seats. Because if you ranked every role by "message density," product managers would almost certainly sit near the top. Picture a typical PM's day: a morning requirements review, where you translate the boss's one-liner "we want an AI feature" into a PRD; an afternoon of cross-team alignment, relaying design's ideas to engineering and engineering's worries back to ops; an evening of chasing progress, updating the schedule, prepping tomorrow's report. How much of that did you create with your own hands, and how much of it was you shuttling information between groups of people, clearing up misunderstandings, syncing the pace? For a lot of PMs, the latter is the bulk of the job. This isn't to say that message-passing has no value. In the past, when information moved poorly inside companies and the walls between departments were thick, a person who could pass the word cleanly and keep everyone in step was genuinely worth money. **The catch is that this value rested on one premise — that moving information was expensive — and AI is driving the cost of moving information toward zero.** Requirements can be auto-structured, progress can be auto-synced, and the cross-team information gap can be flattened by a single shared model. When passing messages becomes free, the role whose whole job is passing messages ends up in the most awkward spot in the room. Product managers aren't going away. But a PM who has bet their entire value on "I can get everyone coordinated" is standing exactly where AI sweeps first this round. ## So which kind of PM survives The contrast makes it obvious. Say two PMs both spot that "this flow is broken." The first one's move: write a doc, schedule a review, relay the requirement to engineering, wait two weeks, get a demo back, take it and check whether it's right. The whole chain is message-passing and waiting. The second one's move: in a single afternoon, use AI to stand up a clickable, runnable prototype themselves, put it in front of users to watch the reaction, and — once it's right — hand it to engineering to harden into a product. Those two weeks of "relay, schedule, wait" got swallowed in one go by doing it themselves. The difference isn't who's smarter. It's that **the first person's value grows out of "information routes through my hands," and the second person's value grows out of "results come out of my hands."** AI can replace the former; it can't replace the latter. Those 7,000 people Meta moved from old roles into the AI division were moved in exactly this direction — from "the message-passing seat that maintains old processes" to "the doer's seat that gets AI to actually produce things." The company doesn't want fewer people. It wants fewer of the seats whose only job is passing messages. So the signal this round of layoffs sends product managers is very concrete: **the better you are at turning "I have an idea" straight into "here's something that works," the fewer messages need to pass through you, and the harder you are to catch in this sweep.** Knowing how to use AI is just the entry ticket. What actually opens up the gap is whether you can turn a single sentence into a running result — by yourself. ## Last thing Microsoft's line — "not replaced by AI" — may be the most honest and most easily overlooked thing said in the past six months. The unspoken second half is this: if a role's entire value is shuttling between pieces of information, then who replaces it is only a question of timing and naming. This year it's called AI. Last year it was called cost-cutting. The year before that, flattening. The name changes every year; the seat that gets picked doesn't. The layoff list gets refreshed every year. The people sitting in the message-passing seat are always near the front of it. ## Further reading - [Every major tech layoff in 2026 that has name-checked AI (TechCrunch)](https://techcrunch.com/2026/07/06/the-running-list-major-tech-layoffs-in-2026-where-employers-cited-ai/) - [2026 Tech Layoffs Near 150,000 as Companies Pour Money Into AI (Yahoo Finance)](https://finance.yahoo.com/sectors/technology/articles/2026-tech-layoffs-near-150-110000224.html) - [2026 tech company layoffs (InformationWeek)](https://www.informationweek.com/it-staffing-careers/2026-tech-company-layoffs) --- # The 100 PMs Who Changed the World · No. 13 | Jensen Huang: The Most Expensive Company on Earth — Stuck Outside the Pantheon of Product Managers URL: https://doaipm.com/en/blog/jensen-huang-sells-the-shovels/ Published: 2026-07-21 Tags: Jensen Huang, NVIDIA, 100 PMs Who Changed the World, Product Management, Tech Commentary Take a look at this month's noise first. OpenAI shipped the GPT-5.6 trio, xAI shipped Grok 4.5, Moonshot open-sourced the 2.8-trillion-parameter Kimi K3, Meta threw out Muse Spark, and even Mira Murati, who left OpenAI, unveiled a model of her own. Every one of them is shouting that they're the strongest; every release is fighting for the headline. The model war is being fought bloody. But glance at the bottom of this melee and you'll notice something a little funny: **these mutually slaughtering models almost all run on the same man's chips.** That man is Jensen Huang. No matter who wins or loses up top, his NVIDIA is selling the shovels — and this month, his market cap soared to $5.4 trillion, the most valuable company on this planet. A man standing at the most central, most guaranteed-to-profit position in the entire AI era. And yet when I had [Claude score the "100 product managers who changed the world"](/en/rankings/), Jensen Huang came in at No. 13, overall OVR 94 — **stuck outside the Pantheon (the 95 line) by one full point.** Spread his six dimensions out and they look like this: **Vision 99 · Insight 87 · Taste 83 · Business 95 · Scale 96 · Originality 95.** Vision 99 is the highest tier in the room; and Taste 83, Insight 87 are the two clearly on the low side of his six. Why can't a boss whose company is worth more than any other get into the Pantheon of product managers? The answer lies in the gap within this set of scores. ## Vision 99: the bet he made at 30 that won him the entire AI era Start with his strongest dimension — the one that most deserves to be remembered. In 1993, Jensen Huang was 30 when he founded NVIDIA to make graphics cards. For a long time afterward, NVIDIA was simply "a graphics-card company for gaming" — that was the full extent of most people's understanding of it. But Jensen Huang bet on something almost no one believed at the time: **the GPU isn't just for rendering game graphics, it's a general-purpose computing platform.** In 2006 he launched CUDA, letting developers use the GPU for general computation beyond graphics. That decision looked nearly obsessive at the time — it added enormous cost, the market simply didn't need it, and Wall Street cursed him year after year for neglecting the core business. CUDA waited, alone, for the better part of a decade. Then deep learning arrived. People discovered that training neural networks required exactly the kind of massive parallel computing the GPU was built for, and the only thing in the whole world that was ready was NVIDIA and the CUDA ecosystem it had toughed out for ten years. **Betting on a future that wouldn't pay off for ten years, and toughing it out through the whole world's incomprehension until it did — that's the caliber of Vision 99.** The people on this list who can earn a 99 on vision can be counted on one hand, and Jensen Huang is one of them. ## Originality 95 and Scale 96: he turned "selling shovels" into the bedrock of an entire era What Jensen Huang truly originated isn't some faster graphics card — it's a position. Before him, a chip company was a "supplier" upstream in the industry chain; after him, NVIDIA became the **bedrock** the entire AI era can't get around. Today, if you want to train any large model — GPT, Gemini, Kimi, whatever — the first thing you do is buy his cards and use his CUDA. He turned NVIDIA from "the guy selling parts" into "the guy defining how the whole game is played." That's the highest form of "selling shovels." In a gold rush, the ones digging for gold win or lose, but the shovel-seller profits for sure — and Jensen Huang doesn't just sell shovels, he monopolized the shovel and, along the way, defined what a shovel even is. At GTC he set a target for two generations of AI chips: by the end of 2027, those two series alone should hit at least $1 trillion in revenue. Scale 96, Originality 95 — no overpayment there. ## Taste 83: this dimension being his lowest is exactly what tells you which kind of product manager he is Now we reach that key gap. Of the six, Jensen Huang's lowest is taste, 83. What is taste? For Jobs, it was "this rounded corner is two pixels off and I simply cannot stand it." For Allen Zhang, it was "in ten years of WeChat, what I'm proudest of is what we didn't build." The core of taste is a person taking their own aesthetics and using them to decide, on the end user's behalf, "what is good." Jensen Huang's products don't face that layer of "user" at all. His customers are developers, data centers, other companies. Whether a GPU is good is measured by compute, efficiency, ecosystem — not rounded corners, color schemes, feel in the hand. **He makes the thing that lets others make good products, not the good product itself.** On the yardstick of "holding the line on aesthetics for the end user," he's naturally out of the running, and Taste 83 is the reasonable result of this position. So that 83 doesn't mean he's weak in ability — it means he simply isn't on that track. **Jobs and Allen Zhang are taste-type product managers, defining a consumer product through personal aesthetics; Jensen Huang is a platform-type product manager, using vision and technical judgment to build the foundation every other product can't do without.** Both kinds are great, but only the former earns taste points. That, together with Insight 87, is why he's held outside the Pantheon door — in that top tier, nearly everyone has maxed out on "defining what is good on the user's behalf," and that's precisely not Jensen Huang's battlefield. ## So does he even count as a product manager? Writing this far, one question is unavoidable: Jensen Huang has never made a single consumer-facing product — does he count as a product manager? By the traditional definition, barely. But this list has never been about titles; it's about **the degree to which product decisions changed the world.** Betting the GPU as a computing platform, nurturing the CUDA ecosystem through a decade of doubt, re-architecting the company from a chip vendor into an "AI factory" — the weight of these decisions is no lighter than that of anyone who made a blockbuster app. He isn't making one particular product; he's deciding what foundation every product of this era grows on. Giving him a 94 and placing him in the Legends tier rather than the Pantheon isn't a put-down — it's precision: **he tops out on "vision" and "originality," and is naturally untouched by "taste," the consumer-grade sense of it.** The score itself says clearly who he is. ## In closing This month's model melee will keep going. GPT-5.6, Grok, Kimi — who ultimately wins, no one dares say right now. But one thing is already certain: no matter who wins up top, Jensen Huang wins. Because what he stands on isn't the ring — it's the ground beneath the ring. A man whose taste scores only 83 has become the most indispensable person in this era that prizes "smart" and "strongest" above all. That points to another path often overlooked: when everyone wants to make the most dazzling product, there's another way to win — to build the foundation every dazzling product can't do without. It doesn't require you to have great taste; it requires the vision to see ten years out, and the endurance to tough out those ten years. Jensen Huang bet right, and then the whole era had to come buy shovels from him. --- *The "100 Product Managers Who Changed the World" list and its six-dimension scoring in this piece were all produced by Claude (AI), rating products and business decisions, not the individuals. For the full list, see [doaipm.com/zh/rankings](/zh/rankings/).* --- # Spain Beat Argentina 1-0 and Lifted the Cup, Messi Bowed Out With 0 Shots: What Locked Down the World's Best Was a System URL: https://doaipm.com/en/blog/the-system-beat-the-genius/ Published: 2026-07-20 Tags: Messi, World Cup, System vs Individual, Product Managers, Tech Commentary Start with the image that will be remembered for a long time. July 19, MetLife Stadium, New York. Lionel Messi, 39, in the final of the last World Cup of his career, played the full match, total shots on the night: **0**. The entire Argentina team took just 2 shots in 120 minutes, not one on target. In the 106th minute of extra time, Ferran Torres broke through one-on-one for the winner, 1-0, and lifted the World Cup trophy. This was not an evenly matched final. The shot count ended up 20 to 2. Argentina's goalkeeper Martínez made 11 saves — the most in the history of a World Cup final — a keeper holding up half the match almost single-handedly through saves alone. And at the other end of the pitch, the player the whole world agrees is the best on this planet, all night, never got to touch a single real chance. The strongest individual on Earth was completely locked down. The question is: who locked him down? ## Spain didn't go find a stronger Messi If you thought Spain shut Messi down by putting some genius one-on-one on him, you misread this match. The way Spain locked Messi down was to **make sure he simply couldn't get the ball**. Rodri anchored the midfield, choking off every passing lane inch by inch; with close to 70% possession, Spain kept the ball at their own feet and left Argentina running without it most of the time — chasing, defending. However strong Messi is, you have to give him the ball first before he can create. And the first thing this system did was keep the ball from reaching his feet. This is the way Spain has won consistently for years: **positional play** — eleven players like a precision machine, each in a specific spot, doing specific passing and specific runs, winning through collective control rather than one person's flash of inspiration. Spain has stars too: Rodri is a Ballon d'Or-caliber midfielder, Lamine Yamal is the next-generation prodigy. But Spain never stakes a match on any one person. Swap out any single player, and the system keeps turning. Here I owe Messi a fair word. Making this man go a whole night without a single shot was never something one defender's man-marking could do — for nearly two decades, teams around the world tried every method, and the overwhelming majority failed. Spain pulled it off this time using the combined force of eleven players, an entire system. **You can count on one hand the players in the world who could force defending at that level.** Messi lost this final, but the very fact that "it takes an entire team, an entire system, to lock him down" already says who he is. Argentina's way of winning over these two decades is a different one. **Give the ball to Messi, and leave the rest to the genius.** That style is unstoppable when Messi is on form — that's exactly how they won it all in 2022. But it has one fatal fragility: when the opponent uses an entire system to lock Messi down, Argentina has no second plan. The whole team with 0 shots on target is the most naked exposure of that fragility. On one side, "hand everything to one irreplaceable genius"; on the other, "hand the outcome to a system that depends on no single person." This final was a head-on collision of those two philosophies. The system won. ## After watching this match, as someone who builds products, I couldn't help thinking of the same thing You might feel it's just a football match — what does it have to do with building products and leading teams. Honestly, at first I was just watching it as a game. But these two ways of winning — the "Messi model" and the "Spain model" — are so much like the choice every team and every product makes each day that it's hard not to draw the connection. Too many teams run the "Messi model" — staking the whole operation on one irreplaceable hero. A star engineer who can handle anything, a genius boss brimming with ideas, a firefighting captain who gets thrown at every crisis. When that person is around, the team is unstoppable, the efficiency is frightening. But the moment that person gets locked down — tangled up in something else, poached away, in a form slump, or just having a bad day — the whole team, like Argentina, puts up 0 shots on target. All your certainty is tied to one variable you can't control. The "Spain model," meanwhile, is about settling capability into a system. **It's not that you don't want geniuses — it's that you don't stake the outcome on a genius's flash of inspiration.** You write "how to get it right" into a process, turn it into a method, harden it into tools, so that every ordinary person standing inside that system can perform at close to the same level. The hero can walk off the pitch, and the system keeps running. This isn't a new idea, but this final acted it out with unusual cruelty and unusual clarity. Look back at the companies that truly grew huge, and almost all of them run the "Spain model." Zhang Yiming took "making a hit" — the thing that by nature depends most on a genius's intuition — and turned it into a machine that can mass-produce; ByteDance got called the "App Factory" — and factory means precisely that it doesn't depend on one genius calling the shots by gut every day. Bezos forced tens of thousands of people at Amazon to write six-page memos, to write the press release first, turning "thinking it through" from one person's talent into a move that tens of thousands can follow. What they did was the same thing: **move the outcome from "whether some particular person is up to it" to "whether this system is up to it."** ## In the AI era, this lesson only weighs more Bring it to today, and the weight of this keeps growing. Because AI is rapidly turning "the strongest execution ability" into something anyone can buy. The strongest models are going open source, dropping in price; what was one company's exclusive capability yesterday is everywhere today. When "capability" itself increasingly becomes like water and electricity, a cruel question surfaces: if the strongest capability is available to everyone, what exactly do you win on? The answer, more and more, is not "we have a genius no one else has" but "we have a system no one else has" — a system that settles judgment, method, and hard-won lessons into place so that every ordinary person can stand on top of it and perform at a high level. There's a point here that's easy to overlook. Judgment is of course precious — I was saying exactly that in my last piece. But if judgment only lives inside one person's head, it is your Messi — precious, and fragile. That person is here, you win; that person gets locked down, and like Argentina you hand in a blank sheet. The truly stable approach is to make judgment part of the system too: write it down, pass it on, turn it into something every person on the team can call up. Individual judgment is genius; systematized judgment is the moat. The reason Spain's title is worth a note from everyone who builds things is right here. They proved something a little counterintuitive in this era that worships the "super individual": **on the biggest stage, a well-designed system can leave the strongest individual on Earth with nowhere to strike.** ## Finally Messi bowed out. 39 years old, his last World Cup, played the full match, 0 shots, the title-defense dream shattered. This one night won't change anything. He is still one of the greatest players in the history of the sport, still the strongest individual. Over these years, how many people started watching football, fell in love with the game, because of him; one man carrying an eleven-a-side team sport on his shoulders — carrying it for so many years, carrying it to nearly forty and still standing on the grass of a World Cup final — is itself close to a miracle. He deserves a standing ovation, in victory and in defeat alike. But what's left on the scoreboard is not only the regret of a genius who couldn't complete the story. There's also a colder fact, one more worth chewing on for everyone who builds things: what locked down the strongest individual on Earth was not another, stronger genius — it was a system. Football will remember that night, and it will remember Messi forever. And those who build products, lead teams, and make things in the AI era had best also remember the method Spain used to lock him down — because the operation in your own hands will, sooner or later, have to make its own choice between "betting on a Messi" and "building a system." --- # Why Does AI Keep Getting It Wrong? It's Not Dumb — You Didn't Say It Clearly URL: https://doaipm.com/en/blog/you-didnt-say-it-clearly/ Published: 2026-07-20 Tags: Describing Requirements, Prompting, Speak It Into Being, Product Managers, AI Workflow Let me start with a scene you've almost certainly lived through. You ask AI to build something, wait eagerly for it to deliver, and what comes back is a universe away from what you had in mind. You fix one line, it breaks another, back and forth a few rounds, until you finally get fed up and mutter to yourself: "This AI is nothing special." These past few months I've been using AI to get work done nearly every day, and I've fallen into this pit every day too. But after enough of it, I changed my mind. Because I noticed a pattern that stings a little: **the same requirement, phrased a different way, and it gets it right the first time.** Once or twice is a coincidence. Eight or nine times is not. Slowly I had to admit something: **when AI gets it wrong, most of the time it isn't dumb — I didn't say it clearly.** This isn't mysticism. Underneath it is a very plain truth, plain to the point of being a little cruel: > AI can't read minds. It only does what you "said," not what you "meant." The requirement in your head is three-dimensional — it has context, it has default assumptions you're not even aware of, it has a whole set of "well, obviously" common sense. But the sentence you type into the chat box is often just a thin outer layer of skin. That layer of skin is all AI gets; the rest it has to guess. Guess right and you got lucky; guess wrong and you go and call it stupid. Once I understood this, I stopped fussing over "which AI is smarter" and turned to practicing something far more valuable: **how to move that three-dimensional requirement in my head, complete, into the chat box.** This piece is that whole method. All of it is copy-pasteable — not one line of it is some talent-only secret technique. And I have to say, this is fantastic news for product managers. **Describing requirements has always been a PM's bread and butter.** You used to explain the requirement to engineers; now you explain it to AI — the only difference is that engineers use experience to fill in what you left out, and they come back to ask you, while AI is more obedient and more "literal," doing exactly as much as you say. So with AI, you have to say it a little more clearly than you would to an engineer. That little bit is the threshold, and it's the dividing line. ## A requirement that lets AI get it right the first time has five parts I pulled apart the requirements I'd "said clearly" and found they always had five things in place. Leave one out, and that's exactly where AI decides for me. You don't have to write all five every time and turn it into a formula. But you need those five slots in your head — **before you hit send, scan them: whichever slot is empty, that's where AI will improvise.** ### Part one: what you want — nail the noun first The most basic, and the easiest to skate past. You say "help me build a user-management thing" — what is this "thing"? A page? A table? A whole back office? AI can only pick one and guess. Change it to: "Build a user list page, one table, showing avatar, name, email, sign-up date, and status." — now it knows exactly where to start. **The whole trick is one line: swap vague words like "thing," "feature," "module" for concrete nouns.** Page, table, button, form, chart, modal. Once you can say what it actually *is*, AI can catch it. ### Part two: for whom, and why — the part most often skipped, and the most valuable This is the part people from a technical background tend to skip, and it happens to be the PM's home turf — you're wired to care about "who uses it and why." Compare. You say "build a dashboard, put all kinds of data on it," and it hands you a screen full of charts you probably won't even look at. Change it to: "Build a dashboard for the **store manager** to check every morning. What he cares about most is **how much sold yesterday, up or down versus the day before, and which category sells best** — put those three front and center, everything else secondary." See the difference? Add "who uses it + what they care about most," and AI knows **what to emphasize and what to dial down** — it starts having "priorities," instead of laying all the information out flat on the floor. I've basically made it a habit now: **after every requirement, I tack on one line about "who this is for, and what problem they're trying to solve."** That one line is often worth more than everything I said before it combined. ### Part three: the concrete form — say what it "looks like" The idea has a picture to it in your head, but all you tossed out was an abstract word. "Build a filter feature" — AI doesn't know what you want to filter by or how. "Above the table, add a row of filters: a date-range picker, a 'status' dropdown (All / Active / Disabled), and a search box (search by name or email); filtering is instant — it refreshes the moment you pick, no confirm button." — now it can reproduce it one-to-one. You don't need to know design. You just need to **describe the picture in your head in plain words**: which blocks there are, what each one is, where it goes, how it's used. Describe it and it can reproduce it; can't describe it, and that means you haven't thought it through yourself — which is perfect, because that's a signal, and later I'll cover how to use it to think things through in reverse. ### Part four: the boundaries — the dividing line between "looks like it runs" and "actually works" I need to spend a little longer on this one, because it's **the real difference between a beginner and a PM.** Beginners describe only the "normal case": the user fills it in obediently, the data is all clean, the network never drops. But in the real world, users type nonsense, data comes back empty, the network cuts out, inputs run too long. These are the "edge cases," and they're what a PM should worry about most — and what you most need to spell out for AI. You say "build a feedback form with name, email, content, and a submit button," and it hands you the happy path and calls it done. If instead you say this: > "...watch out for these cases: when the email format is wrong, show a red hint under the field and don't let it submit; when the content is empty, the submit button stays disabled; on success, clear the form and show 'Thanks for the feedback'; on failure (say, no network), **don't clear** what the user already typed — show 'Submission failed, please try again.'" Those few extra lines are the PM's professionalism. AI is fully capable of handling all of this, but if you don't say it, it defaults to just the "happy path." List the boundaries, and it builds them all in the first time. By the way, **"what counts as done" and "what not to build" are boundaries too, and just as worth saying.** A single line — "this version does only the list and the filters, no create/edit/delete yet" — blocks a whole pile of things you didn't want this round. ### Part five: a reference — hand it a ruler Style, colors, feel — these "vibe" things are the hardest to pin down in words. The easiest move: give it a reference. "Make the colors nice, make it look professional" — a hundred people have a hundred readings of "nice" and "professional." "Use brand blue #3B82F6 as the primary color; for the overall palette and whitespace, follow the clean, restrained style of the Stripe website." — now it has a ruler and doesn't have to guess at the "nice" in your head. A reference can be a color value, a product you like ("like Notion"), a standard. Give it a reference, and AI doesn't have to bet on your taste. ## Five most common ways of "not saying it clearly," and how to fix each The above was the constructive breakdown. Below are the five pits I see most, and fall into most myself — go down them one by one as a checklist. **Pit one: too vague.** "Build a nice-looking page" — it hands you something mediocre and you can't even say what's wrong. Fix: push it one level more concrete, swap adjectives for nouns and details. **Pit two: giving orders with no context.** "Add an export feature" — export what? Export to where? What format? Add it where? Fix: fill in "where + what to export + what format," e.g. "on the reports page, add an 'Export' button top-right; clicking it exports the table under the current filters as a CSV download." **Pit three: dumping a whole pile at once.** Cram fifteen requirements into one message, and what comes back is a mess you don't even know where to start fixing. Fix: **small steps, fast — one at a time.** First put up the skeleton, run it, see the result, then add on one by one. This is the most important of them all, no contest. **Pit four: no reference, leaving it to guess your taste.** Fix: see part five — give a color value, a reference product, a standard. **Pit five: only saying the normal case, not the abnormal.** The demo runs smooth, then real users touch it and everything breaks. Fix: before you send the requirement, ask yourself three questions — **what happens if the data is empty? what happens if the user types nonsense? what happens if the network drops?** Write the answers into the requirement. ## Three moves: turning AI from a hand into a brain Get the above down and you're already ahead of most people. The three moves below are what I use to turn AI from "an obedient pair of hands" into "a brain that helps me think." **Move one: have it ask you questions first, don't rush it into building.** This is the one I use most, and the most counterintuitive: the more important the requirement, the less you should have it build right away — have it question you first. > "I have an idea: build a tool to collect and analyze customer feedback. Don't start yet — first ask me the five most critical questions, and pin down the target user, the core use, the must-have features, and the boundaries." The questions it asks are, nine times out of ten, exactly the spots you haven't thought through yourself. The process of answering is the process of forcing the requirement from "a fuzzy blob" into "a clear line." By the time you're done answering, it's holding a complete requirement, so of course what it builds is right. The essence of this move is **using AI to help you think clearly, not just to help you build** — think it through, and getting it right comes for free. **Move two: give a good example and a bad one.** When words fall short, an example is fastest — especially "I want this kind, not that kind." > "Help me write the message that shows after this button is clicked. Make it like this: short, casual, reassuring — 'Saved, you're good.' Not like this: stiff, wordy — 'Your operation has been successfully submitted to the server and persisted to storage.'" One good example plus one bad one beats three lines of adjectives. **Move three: have it restate to confirm, then build.** When the requirement is complex, have it restate its understanding first, and only build once you've confirmed it's right. > "The requirement I just gave you — don't build it yet. Confirm it back to me in three sentences: what you're about to build, how many parts, and anything you're unsure about. Once I confirm, then start." Thirty seconds here saves you half an hour of rework. As it restates, you'll often catch on the spot that "wait, what I said here isn't what I meant" — fix it right then, before it's touched anything. ## One full run-through: from a fuzzy sentence to it getting it right the first time Let's string the above together and watch a real run. Say all I have in my head is one fuzzy sentence: "I want a thing that lets me look at user feedback." Step one, I don't rush it into building — I have it help me think first (move one): I have it ask me five key questions. It asks: who uses it (my ops colleague)? where does the feedback come from (a CSV file)? what do I most want to get out of it (quickly see what everyone complains about most)? do I want it sorted by severity (yes)? do we need filters this version (not yet)? Step two, fold the answers into one complete requirement, and the five parts line up perfectly: for ops to look at, the goal being to quickly see what users complain about most (for whom, why); an upload area up top to load the CSV, and below it auto-display the feedback grouped by theme, how many items per theme, the share of each, sorted by count (what, form); when the CSV is empty or the format is wrong, show "Read failed, please check the format," and no filters or search this version (boundaries); clean and restrained, primary color #3B82F6, layout following Notion (reference). Build it out with fake data first, and spin up a local server so I can preview. Step three, run it, see the result, then add in small steps (the fix for pit three): one at a time — "add a progress bar after each theme to show the share," "clicking a theme expands to show three raw feedback items in that category," "color themes by severity — red for high frequency, yellow for medium, gray for low." Step four, have it find the problems itself: "click through everything clickable, find the errors, hangs, and anything not behaving as expected, and give me a list." The whole way through I didn't write a single line of code, but every step I said clearly. This is the daily reality now — **you're in charge of thinking it through and saying it clearly, it's in charge of building it.** ## Finally Before I hit enter, I usually spend ten seconds running these slots through my head: what I want (is it a concrete noun?), for whom and why, what it looks like, boundaries (empty data, garbage input, failure, what not to build — did I cover them?), reference, and — is this just one requirement this time. Get four or five of the six, and it basically gets it right the first time. More and more people know how to use AI, but the people who can say the requirement clearly have always been the minority. The first thing is a threshold — everyone can cross it; the second is a dividing line — it splits people into two camps. And you, as a product manager, already understand users best, understand priorities best, understand the very boundaries no one else will worry about for you. What you've been missing was never those things — it's the habit of **saying them out loud.** Practice for a month, and you'll find it's not that AI got smarter — it's that you finally learned how to hand over that requirement in your head, complete. --- # The Woman Who Built ChatGPT Just Shipped a Model She Admits Isn't the Strongest — and Investors Are About to Value Her at $50 Billion URL: https://doaipm.com/en/blog/mira-murati-not-the-strongest/ Published: 2026-07-19 Tags: Mira Murati, Thinking Machines, OpenAI, AI, Product Managers, Tech Commentary Start by putting two things side by side. Yesterday, China's Kimi K3 was all over the timelines, and the press-kit keywords were "the world's largest open-source model," "topping the coding charts," "beating GPT and Claude" — everyone was comparing whose scores were higher. Today, another model shipped, and its wording ran the other way. On July 16, Mira Murati's Thinking Machines dropped its first open-source model, Inkling — 975B parameters, natively multimodal, with thinking effort you can dial from 0.2 up to 0.99. But in the official release notes, one line stands out sharply: **"not the strongest overall model available today, open or closed."** A company releasing its flagship model and voluntarily telling you "it isn't the strongest." In an industry that today would happily print every benchmark win onto a poster, that's practically heresy. Stranger still: this very company, just a year and a half old, is raising a new round at roughly a **$50 billion** valuation — more than four times its $12 billion last July. Why is someone who shipped a model that isn't the strongest worth $50 billion? To make sense of that, you first have to know who Mira Murati actually is. ## Plenty of people can train models; almost no one can turn a model into a product Mira Murati spent six and a half years at OpenAI, rising all the way to CTO. But her real weight isn't in those three letters. She was the key operator who took ChatGPT, DALL·E, and GPT-4 from the lab to the public. In late 2022, OpenAI had a powerful model that ordinary people couldn't touch; it was her layer that wrapped it into a chat box anyone could open and use — and then ChatGPT became the fastest-adopted consumer product in human history, a hundred million users in two months. During those days in November 2023 when the board abruptly ousted Altman, the person who stepped in as interim CEO was also her. Put differently: Silicon Valley has no shortage of people who can train models, but the people who can turn a model into a product a billion people want to open every day are a tiny handful. And on that very short list, Mira Murati's name sits near the top. That's the key to understanding her — and to understanding the $50 billion valuation. ## That "not the strongest" line in Inkling is exactly what exposes her playbook Back to that unusual release. Someone who understands better than anyone that "a model is not a product" shipped a model that doesn't chase the top spot — that's a strategy someone thought all the way through. Look at what Inkling leads with and it's clear: native multimodal reasoning (designed for real-time human-machine interaction, not retrofitted from a pure-text model), adjustable thinking effort (get the same job done with fewer tokens), and fine-tuning that works out of the box on its own Tinker platform. It even released a 276B smaller version, Inkling-Small, that matches the big model on several benchmarks at far lower cost. Not one of these selling points is "I rank first on some chart." They all point at the same thing: **can you actually put it to use, wire it in smoothly, and conveniently take it away to build your own thing.** This is the hardest lesson Mira Murati took from ChatGPT — what decides whether an AI product lives or dies is rarely where it ranks in the evals, and much more often whether anyone can hold it in their hands and turn it into something others want to use. Yesterday's Kimi K3 and today's Inkling stand at exactly the two ends of this. One makes "topping the charts" the headline; the other writes "I'm not the strongest" into the release. Two philosophies — it's too early to say who's right — but which side Mira Murati is betting on is already down on paper. ## The $50 billion is a bet on her, the person Look again at that valuation curve and it gets even clearer what the market is pricing. Thinking Machines was founded only in February 2025, and by July that year it closed a $2 billion seed round — led by a16z, with NVIDIA, AMD, Cisco, and Jane Street joining, at a $12 billion valuation. That was already one of the largest seed rounds ever. Today it's raising at $50 billion, a fourfold jump in a year, muscling straight into the ranks of the world's most valuable private companies. Leaving OpenAI alongside her were co-founder John Schulman, former VP of research Barret Zoph, and that whole core crew. Investors aren't fools. They know Inkling isn't the strongest model right now, and they know this company hasn't yet built a breakout product like ChatGPT. Betting at $50 billion, what they're wagering on is Mira Murati the person — that kind of judgment for "turning research into a product a billion people use." She cashed that ability in once before; now they're betting she can cash it in again. ## As models become more of a commodity, judgment is the scarce thing The real signal here is bigger than "another star startup." Look at what's happened these past six months: Kimi gave away a 2.8-trillion-parameter model for free, open-source flagships are everywhere, and the raw capability of models is getting cheaper and more readily available at a pace you can watch with the naked eye. When capability itself becomes as commoditized as water and electricity, a brutal question surfaces: if everyone can download the strongest model, then what is actually worth anything anymore? Mira Murati's $50 billion valuation is one answer the market is giving: **people who can train models are no longer scarce; people who can judge "what is good, which direction to push, and how to package a capability into a product people want to use" — those are scarce.** The former can be stacked up out of open source and compute; the latter can't. She's worth it because she knows what a good product is, not because she can build the strongest model. ## This $50 billion bet has only just begun That said, Inkling really isn't the strongest model right now, and Thinking Machines hasn't yet proven it can pull off a ChatGPT-scale miracle a second time. This $50 billion bet is nowhere near being settled. But one thing has already been written into this release: in an era where models look more and more like a commodity, someone who understands better than anyone that "scores are not a product" shipped a model that doesn't chase scores — and then landed the most expensive vote of confidence in the whole industry. Whether Inkling is the strongest doesn't matter; what matters is that the thing she's betting on — judgment is worth more than capability — is being voted for by more and more money. ## Further reading - [Inkling: Our open-weights model](https://thinkingmachines.ai/news/introducing-inkling/) — Official release from Thinking Machines Lab - [Murati's Thinking Machines in Funding Talks at $50 Billion Value](https://www.bloomberg.com/news/articles/2025-11-13/murati-s-thinking-machines-in-funding-talks-at-50-billion-value) — Bloomberg - [Thinking Machines open sources Inkling, focused on low cost and 'resistance to censorship'](https://venturebeat.com/technology/thinking-machines-open-sources-first-multimodal-language-model-inkling-focused-on-low-cost-and-resistance-to-censorship) — VentureBeat --- # The 100 PMs Who Changed the World · No. 3 | Jeff Bezos: At 62 He Made Himself CEO Again — to Bet on the One Thing Amazon Never Pulled Off URL: https://doaipm.com/en/blog/bezos-working-backwards/ Published: 2026-07-18 Tags: Jeff Bezos, Amazon, AWS, Product Managers, 100 PMs Who Changed the World, Tech Commentary Start with something that happened just last month that many people didn't much notice. Jeff Bezos, 62, handed off the Amazon CEO job back in 2021, and for the four years since, the public image of what he was doing ran roughly like this: building rockets, sailing yachts, holding a wedding in Venice. Bloomberg pegs his net worth at around $250 billion, top three in the world. A man like that should, by all conventional logic, have entered the "made it, stepped back, just spend the money" phase of life. Instead, in 2026, he made himself CEO again — not back at Amazon, but personally co-leading an [AI startup called Project Prometheus](https://www.stcn.com/article/detail/3500750.html), aimed at "AI for the physical world": putting large models to work in the engineering of manufacturing, automobiles, and aerospace. The company just closed a $1.2 billion round in June at a $41 billion valuation. Add in his money in Anthropic and others, and the capital he's staked on AI already tops $19 billion. That's a little strange. **Bezos is the world's most skilled product manager at "letting go."** The Amazon he built by hand is defined, above all, by not depending on him — the flywheel spins on its own, the systems run on their own, and he retreated backstage early. So why would a man who perfected "make the company run even without me" find himself, on this one thing called AI, unable to resist sitting back down in the CEO chair at 62 and stepping into the ring himself? To think that question through, you first have to understand what he actually invented. When I had [Claude score the 100 product managers who changed the world](/en/rankings/), Bezos came in at No. 3, overall OVR 97. Spread his six dimensions out and it looks like this: **Vision 98 · Insight 95 · Taste 90 · Business 99 · Scale 98 · Originality 97.** That Business 99 is the highest tier on the whole list (tied with Bill Gates), one point above Jobs's 97. This piece starts from that 99 — because it hides a misconception most people hold about Bezos. ## Business 99: he turned "doing business" itself into a product Mention Bezos and many people's first reaction is "e-commerce," "logistics," "the guy who perfected package delivery." That impression isn't wrong, but it sells Bezos short. What he's truly great at is designing the business model itself as a product. The classic example is AWS. In the mid-2000s, to support its own e-commerce, Amazon was forced to build out a huge stack of servers and operations capability. The vast majority of companies would treat that as a cost, an internal department, and stop there. Bezos made a decision no one understood at the time: **take the infrastructure you built for internal use, package it as a product, and sell it to everyone.** And so the cloud computing industry was born. Today every internet company in the world, including the vast majority of AI startups, runs on someone else's cloud — and Amazon was the one that first blazed that trail. In Q1 2026, AWS alone brought in $37.6 billion in quarterly revenue. Notice what happened here: he didn't invent a new app, didn't build a prettier interface. The product he built was a business structure — a way to "turn cost into revenue, turn internal capability into an outward business." The same thinking runs through Prime membership, the third-party seller marketplace, Kindle — none of them a simple feature, each a full, self-reinforcing business machine. That's why, on the Business dimension, he can pull the list-high 99: when it comes to "designing the act of making money itself as a product," no one on this list did it more thoroughly. And that ability comes from something more fundamental. ## Originality 97: Amazon's real moat is a method that forces people to think clearly Of everything Bezos left behind for Amazon, the most underrated is two writing rules. **Rule one: no PowerPoint in meetings — you submit a six-page narrative memo.** Amazon's important meetings open with everyone silently reading the document, then discussing. The reason is simple: PowerPoint can take an idea that hasn't been thought through and, with a few slick phrases and one chart, dress it up to look like something; but forcing you to write six pages of complete sentences exposes every logical hole, every "I actually haven't figured this out," right there on the page. **Rule two: Working Backwards — write the press release first, then build the product.** Before a new idea gets greenlit, you pretend it's already built and about to launch, and you write the press release and the customer FAQ. If you can't even write clearly, or write compellingly, "what exactly makes this good for the user," then it probably isn't worth building. Put those two rules together and they're the same thing: **before you build anything, you're forced to write down, in clear language, what exactly I want, for whom, and why it's good.** Amazon's decades, its tens of thousands of people, its hundreds of product lines — none of it rides on one genius making every call each day. It rides on a method that lets tens of thousands of people independently think it through, write it down, then execute. That is the true bearing on which the flywheel — the one that "spins even without Bezos" — turns. Originality 97: half of it goes to AWS, half to this method. ## Taste 90: his lowest score of the six is exactly why he could let go Of the six, the only one Bezos didn't clear 95 on is Taste, at 90. That's not to say his aesthetics are poor — it's a different path. Jobs and Allen Zhang bet on "use my personal judgment to define, on the user's behalf, what is good"; Bezos bet on a process — using the six-page memo and Working Backwards to turn "thinking it through" from one genius's intuition into a move that tens of thousands of people can follow. Jobs's products are extensions of his personal taste; without him, the flavor of Apple would change. Bezos's products are the output of a method, which is why he could retreat backstage early and the machine keeps running. Taste 90 is because he deliberately chose not to stake the company on his own aesthetics — that's his weakness, and it's exactly the confidence that let him dare to let go. ## That method landed right on top of the AI era Now back to the question we opened with: why would the man best at letting go come out of retirement to do it himself at 62? Take one look at what's scarcest in the AI era and the answer surfaces. When writing code, doing design, and building prototypes — all the execution steps — can be handed to a model, what's left, the truly hard and truly valuable part, is **saying clearly what exactly you want** — clear enough that the machine can follow it and build precisely the thing you had in mind. Being vague with an AI only gets you a pile of stuff that looks the part and is actually useless; the person who can spell out the requirements, the boundaries, and what counts as good, one by one, is the one who can actually make AI produce something useful. And that is exactly the thing Bezos drilled into Amazon for decades. **The six-page memo and Working Backwards were originally an internal management system, built to keep a big company from getting confused; in a world where everyone can build things with AI, they've become a universal, most-core productivity.** Whoever can write clearly what they want can build it — a line that in 2019 held true only for Amazon's managers, and in 2026 holds true for every single person sitting in front of an AI. So when Bezos comes out to bet on Project Prometheus, what he's really betting on is a rerun of the thing he knows best. Back then AWS was "turn internal capability into the infrastructure of a new industry"; today the "AI for the physical world" he's staking on is a bet of the same shape — AI has already been proven out in software and text, and the next battlefield is real-world engineering: manufacturing, automobiles, aerospace. What those fields lack most is precisely the ability to "break a complex goal down clearly, write it down, and have the system execute it." This is a thing he doesn't trust to anyone else — so he does it himself. ## A real question Whether this bet pays off, no one knows yet. AI for the physical world is far harder than software; Blue Origin's rockets have blown up and [slipped their schedules](https://finance.sina.com.cn/stock/usstock/c/2026-07-08/doc-inihckrq8234720.shtml) over the years, and Prometheus at a $41 billion valuation could well be a very expensive failure. But one thing has already been proven: when Bezos, more than twenty years ago, forced a room full of managers to write six pages and write the press release first, no one thought it was any great invention — they even found it tedious and old-fashioned. Looking back today, that clunky method of "figure out what you want and write it down clearly, first" may be worth more than any single item Amazon ever sold — because it happens to be, in this era, the single most critical interface between human and machine. A man who has used that method his whole life is choosing, at 62, to stake his fortune and reputation on it again, betting it can win one more time in the AI era. And this time the battlefield is the physical world itself — tougher to crack than either retail or the cloud. --- *The "100 Product Managers Who Changed the World" list and its six-dimension scoring in this piece were all produced by Claude (AI), rating products and business decisions, not the individuals. For the full list, see [doaipm.com/zh/rankings](/zh/rankings/).* ## Further reading - [前首富贝索斯出山!投身人工智能](https://www.stcn.com/article/detail/3500750.html) — Securities Times, on Bezos returning as CEO and co-leading Project Prometheus - [Bezos: My time at Amazon, Blue Origin and Prometheus is spent on AI](https://www.cnbc.com/video/2026/05/20/jeff-bezos-my-time-at-amazon-blue-origin-and-prometheus-is-spent-on-ai.html) — CNBC, Bezos in his own words on how his energy is all on AI right now - [蓝色起源首轮外部融资,估值达 1300 亿美元](https://finance.sina.com.cn/stock/usstock/c/2026-07-08/doc-inihckrq8234720.shtml) — Sina Finance --- # The 100 PMs Who Changed the World · No. 5 | Zhang Yiming: His Best Product Isn't Douyin — It's a Machine That Mass-Produces Hits URL: https://doaipm.com/en/blog/zhang-yiming-the-app-factory/ Published: 2026-07-17 Tags: Zhang Yiming, ByteDance, Product Management, 100 PMs Who Changed the World, Tech Commentary Start with a slightly odd comparison. Last month, the Bloomberg Billionaires Index updated, and ByteDance founder Zhang Yiming, with a net worth of $92.8 billion, overtook India's Ambani to become the second-richest man in Asia — while sitting comfortably atop China's own list. Counting from the $13 billion he was tracked at back in 2019, his wealth has grown more than sevenfold in six years. Meanwhile, a leaked ByteDance strategy document boiled down to a single headline: 160 billion yuan into AI in 2026. Its Doubao assistant has already crossed 300 million monthly actives. The odd part is this: a man standing on the second tier of Asia's wealth pyramid almost never gives an interview, stepped down as CEO back in 2021, and shows his face pitifully rarely. You can instantly picture Jobs on stage at a launch event, or every one of Musk's posts on X — but you probably can't recall what Zhang Yiming looks like, or what he said the last time he spoke in public. **How does an almost "invisible" man become the second-richest in Asia?** The answer to that question hides in a single number. When I had [Claude score the 100 product managers who changed the world](/zh/rankings/), Zhang Yiming came in at No. 5, overall OVR 96. But spread his six dimensions out and one number is glaring: **Vision 97 · Insight 95 · Taste 88 · Business 97 · Scale 98 · Originality 96 — overall 96. Taste 88 is the only one of the six that didn't clear 90.** And here's what I want to say: **that lowest-in-the-room taste score isn't his weakness — it's precisely his sharpest weapon.** This piece starts from that 88. ## Originality 96: he didn't build a product, he built a machine that mass-produces hits Start with his most counterintuitive and most underrated dimension — originality. Most people remember Zhang Yiming for Toutiao, Douyin, and TikTok. But if you see him only as "the man who made Douyin," you've sold him short. Because before Douyin there was Toutiao; after Douyin there was TikTok, CapCut, Feishu, Fanqie Novel, and now Doubao — **making one hit in a lifetime is luck; making hit after hit for ten years straight can't possibly be luck. It can only be a repeatable method.** What Zhang Yiming truly originated isn't any single app. It's a **machine that industrializes the act of "making a hit."** You've probably heard of the machine's core parts: a recommendation algorithm so strong it borders on unnatural, taking "you might also like" to its absolute limit; a middle platform that lets new apps reuse capabilities that have already been proven, so a new one can sprout in a matter of months; a thoroughly data-driven culture where even the color of a single button is decided by A/B test, not by someone's gut. **Before him, making a product was a "craft" — riding on the intuition and taste of one genius product manager. After him, making a product could be "industry" — riding on a system, a pile of data, an assembly line.** That's why ByteDance is called, inside and out, the "App Factory." "Factory" is an insult when it's applied to someone else; applied to Zhang Yiming, it's his greatest achievement: he took something that used to depend heavily on individual genius and turned it into something you can mass-produce at scale. No Chinese product person before him had truly made this road work. Originality 96 — no argument there. ## Scale 98: he turned "dopamine" into a global business Scale barely needs arguing. Douyin plus TikTok add up to monthly actives in the billions — among the most frequently opened apps on the planet. ByteDance's 2025 net profit reportedly reached $48 billion, averaging more than $660 million earned every single day. **What matters more is that this scale is global** — TikTok is one of the very few Chinese internet products to truly take root among mainstream users in the U.S., something not even Tencent or Alibaba managed. But let me pause here and say something not so flattering. The hits this machine of his mass-produces run, overwhelmingly, on the same fuel: **your attention, and that little bit of dopamine you find so hard to control.** Douyin's recommendation algorithm is strong not because it understands "what you need," but because it understands, all too well, "what will keep you from stopping." That's the truth behind Scale 98 — the one we all quietly know and don't much like saying out loud. And it leads straight into his next score. ## Insight 95: he doesn't see needs, he sees the weak spots in human nature When we traditionally praise a product manager for "strong insight," we mean they can see through to the real needs a user never voiced. Zhang Yiming's insight is a different kind — colder, and more effective. **What he sees isn't "what people want," it's "what people can't help clicking."** The difference between those two is subtle but enormous. The former cares about your goals; the latter cares about your soft spots. One treats your intention to become better as its North Star; the other treats your instinct to get distracted, to be curious, to scroll one more, and one more, as fuel. I don't think that's an insult. **Seeing the weak spots of human nature this clearly — and engineering them into an algorithm — is itself top-tier insight,** worth a 95. It's only once you see clearly what this insight is aimed at that you understand why this list can only give him an 88 on the next dimension: taste. ## Taste 88: it's not that he has no taste — he deliberately gave it up Now we reach the key to the whole piece. What is taste? For Jobs, it was "this rounded corner is two pixels off and I simply cannot stand it." For Zhang Xiaolong, it was "in ten years of WeChat, what I'm proudest of is what we didn't build." The core of taste is a person taking their own aesthetics and judgment and using them to decide, on the user's behalf, "what is good" — even when the data doesn't support it yet, I still believe this way is more right. And Zhang Yiming is **precisely the man who deliberately deletes that "I believe."** He has a widely circulated line, roughly: don't use your value judgment to make choices for users — let the data speak. It's an extremely clear-eyed and extremely effective product philosophy — it freed ByteDance from the randomness of "the boss decides on a whim," made every decision traceable, optimizable, scalable. Behind every Douyin redesign, every recommendation, isn't one person's taste — it's the result of tens of millions of A/B tests. But turn that line around and listen again: **the flip side of "let the data speak" is "I won't judge what's good for you."** Data only tells you "which one the user clicked," never "which one is better for the user." It can tell you "the button that gets you to scroll ten more minutes of short video is easier to click," but it will never answer, on your behalf, "whether those ten extra minutes were good or bad for this person." Only taste can answer that question — and in Zhang Yiming's system, that question is precisely the one switched off. So Taste 88 doesn't mean his aesthetics are poor or his ability lacking. It means this machine of his is designed, from the root, not to aim at "is it good" but only at "do you love watching it." **He didn't lose on taste — he simply never put taste into the machine. That's his strongest point, and it's exactly why this list can only give him an 88.** A man who hands "what is good" over to the data will, on a dimension that measures "how much you held the line on good for your users," necessarily fall short of full marks. Set them side by side and it's clearer still: Jobs's products are extensions of his taste; Zhang Yiming's products are what the data grows into on its own, once he's deliberately pulled taste out. On this list, the two men sit at opposite poles of product philosophy. ## Two men named Zhang: one wants you to leave, one wants you to stay And there's a contrast on this list that cuts deeper than Jobs, because he too is Chinese, and he too ranks above Zhang Yiming — No. 2, Zhang Xiaolong. Zhang Xiaolong has a line quoted countless times: **"A good product should let the user use it and leave."** In ten years of WeChat, what he's proudest of is "what we didn't build" — no splash-screen ads, no read receipts, no piling on a bunch of things you didn't want while you're mid-chat. His entire product philosophy is **shielding users from the designs that would pull them into addiction while doing them no good.** Zhang Yiming's machine does almost the opposite: **by every possible means, keep you from leaving, get you to stay a little longer, get you to scroll one more.** Every "next video autoplays" and "infinite feed of things you might like" on Douyin optimizes the same metric — user time on app. Use it and leave? That's the last thing Douyin wants to see. I have no intention of ruling one above the other here. Both made products that changed the daily lives of a billion people; both sit firmly in this list's top five. I only want you to see this astonishing fact: **the question of "what makes a good product" got completely opposite answers from two Chinese product managers who both reached the very top.** One holds that a good product "lets you leave and doesn't waste your time"; the other holds that a good product is "impossible to put down, one more second is one more second gained." And the market, at a scale of a billion, rewarded them both at once. That's also why, on taste, Zhang Xiaolong scores higher and Zhang Yiming lands at 88 — not a gap in ability, but the fact that the two of them stand at opposite ends on "whether to hold the line for the user." Zhang Xiaolong chose to hold, so there are things his products refuse to do; Zhang Yiming chose to let go, handing the judgment entirely to the data and to human nature, so his machine runs faster than anyone's — and cares less than anyone's about whether you feel empty or fulfilled after you've scrolled. ## Vision 97 and Business 97: an "invisible" man in actual control Two dimensions left, quickly. Business 97 needs little explanation — turning attention into advertising, e-commerce, short dramas, and AI across the board, building a money printer that earns $660 million a day. Commercially, he's nearly flawless. On Vision 97 I want to flag an easily overlooked detail: Zhang Yiming stepped down as CEO back in 2021, saying publicly he'd shift to "long-term strategy and company culture." Many took this as retirement. But look at the equity structure and you'll see he holds more than 60% of ByteDance's voting rights — he's the real party in control. **He didn't leave the table; he went from "the man playing each hand" to "the man setting the table's rules."** A person who, at the peak of his career, deliberately retreats from the spotlight into the wings while keeping a firm grip on the final decision — that clarity about "where I ought to sit" is itself a form of vision. ## That machine is now being retested by Doubao Writing this far, the question from the opening — "how does an invisible man become the second-richest in Asia" — has a clear answer: because what he built isn't a product that needs him on stage, it's a machine that needs no appearance from him and keeps mass-producing hits on its own. The more quietly the machine runs, the more invisible he can be. But now that machine faces its biggest test yet: AI. ByteDance plans to pour 160 billion yuan into AI in 2026, with Doubao's monthly actives charging to 300 million. On the surface, this is just the "App Factory" mass-producing its next hit. But AI has one fundamental difference from short video: **the winning move in short video is "who understands better how to keep you from stopping," while the winning move in AI is likely "who understands better what the right answer is" — and the latter is exactly what needs taste and judgment, precisely the thing his machine deliberately pulled out years ago.** So Zhang Yiming's real bet isn't whether Doubao can stack up another 300 million monthly actives — with his machine, that's not hard. The real question is: **can a machine designed to "not judge good from bad, only chase what you love watching" do well at something that "must judge good from bad"?** Can a man who left his taste at 88 make up those 12 points on a new track that rewards taste? That answer will emerge slowly as this 160 billion goes in during 2026. I don't have it yet either — but I do know that this time, "let the data speak," on its own, may not be enough. --- *The "100 Product Managers Who Changed the World" list and its six-dimension scoring in this piece were all produced by Claude (AI), rating products and business decisions, not the individuals. For the full list, see [doaipm.com/zh/rankings](/zh/rankings/).* --- # I handed AI about half of my PM job, and there are a few things I still don't dare hand over URL: https://doaipm.com/en/blog/pm-what-i-gave-ai-what-i-kept/ Published: 2026-07-16 Tags: product manager, AI workflow, judgment, career, division of labor, hands-on Let me start with something that might feel a little counterintuitive: **this year I handed AI roughly half of my day-to-day PM work, and I handed it over without a second thought.** But a few other things I haven't dared hand over, and probably never will — not because AI can't do them, but because once those things are handed off and go wrong, I can't catch it, and by the time I do, it's already too late. What I've found separates these two piles has nothing to do with "hard vs. easy," or "boring vs. interesting." It comes down to a different question: **if AI got this wrong, could I catch it on the spot?** If yes, I'll hand it over. If no, I keep a death grip on it. So I'll walk you along that line, laying out the things I "handed over" and the things I "kept" this year, one by one — and I'll tell you about the couple of times I nearly tripped after handing something off. If you're also stuck on which things to let AI do and which not to, maybe it'll save you some trial and error. ## First, the ones I handed over: the work "I can eyeball as right or wrong" **Number one: first drafts of all kinds of documents.** PRDs, weekly reports, requirement specs, updates for the boss — I basically let AI draft all of these first. It gives me a seventy-or-eighty-out-of-a-hundred foundation, and I edit on top of it. Why do I dare hand it over? Because whether a document is right, I know the moment I read it — which sentence is fluff, which point got missed, none of it escapes me. It pulls me out of "staring at a blank page," and I take its seventy up to the ninety I actually want. It does this faster than I would, and I verify it fast too — that's a good trade. Honestly, the most draining part of writing a document was never the typing; it's the stall of starting from zero. Once that stall gets taken off my plate, what I save is mental energy, not just time. **Number two: digging up material, doing the first pass of competitive research.** When I want to get up to speed on a new area, or size up a few competitors, I have AI pull public information into a first version for me: how each of them does it, where their approaches differ, roughly where they're strong and weak. It does in half an hour what would take me a day of hunting. The key here, again, is "I can verify it" — whether what it organized is right or made up, I can tell by checking it against two points I already know well; if it holds up I use it, if not I redo it. One time it invented a detail about a competitor — "they've launched a paid membership tier" — described in convincing, specific terms, except I happened to know that company well and there was no such thing. I casually checked it against two familiar points and it fell apart. So I never trust the material it digs up outright; I always poke at two spots I understand first, and if it cracks there, the whole version goes in the trash. **What I can hand over is never "I trust it," it's "I can catch it any time I want."** **Number three: sorting a big pile of user feedback and finding the common threads.** Hundreds or thousands of comments and pieces of feedback sitting there — reading them one by one will burn your eyes out. Now I just throw them at AI and have it categorize, tag, and pick out the complaints that keep recurring. It gives me a map of "what users are actually griping about and praising." This used to eat most of my day; now it's half an hour, and it's more patient than I am — it won't start zoning out and skipping lines by the two-hundredth comment. **Number four: meeting notes, and turning discussion into action items.** After a meeting, I have AI turn the recording or transcript into notes and pull out who's supposed to do what; I skim it, add a couple of lines, and send it out. That half hour of cleanup after every meeting — gone. And that skim isn't for show either: now and then it'll assign the wrong owner to something, or treat an offhand remark as a firm conclusion, but those I can spot at a glance and fix; it's precisely because it makes those mistakes that I never skip the skim. **And building prototypes.** I wrote about this one specifically last time — a one-sentence idea, and in an afternoon AI turns it into something you can click. That's in the "handed over" pile too. String these five together and you'll notice they share one thing: **each of them has a right-or-wrong I can check on the spot.** Whether a document is good, I can read. Whether the material is real, I can cross-check. Whether the feedback is sorted right, I can spot-check. **AI frees me from the tiring, time-eating grunt work — and that final "is it right" gate is always still mine to hold.** That's the entire basis for my handing it over with peace of mind. ## Now, the ones I kept: the work "I can't tell is wrong even when it is" The more I hand over, the clearer it gets which few things I absolutely cannot. Because they all step on the same landmine: **when AI gets it wrong, you can't tell on the spot — sometimes you never tell at all.** **Number one: deciding whether to do something, and what to do first.** Prioritization and trade-offs, in other words. AI can help me list all the options and lay out the pros and cons of each — that I welcome. But the decision of "these three requirements, which one do we cut, which do we ship first" I never hand to it. Because prioritization has no verifiable, standard right answer — it depends on what we actually want this quarter, which direction we're betting on, what we're willing to give up, and all of that lives in my head and in countless bits of context never written into any document. **The ranking AI gives me always looks reasonable, and "looks reasonable" is exactly the dangerous part — a wrong priority takes months, takes a whole pile of resources going down the drain, before you realize the direction was wrong from the start.** That kind of error — invisible at the time, unrecoverable after the fact — I don't dare outsource. **Number two: judging whether the stuff AI itself gives me is actually right.** This one sounds circular, but it's especially deadly. AI will very confidently present something wrong as if it were true — a number that doesn't exist, a user conclusion it took for granted, a chain of logic that sounds airtight but doesn't hold up. **If I hand off even the "judge whether it's right" part to AI (say, sending in another AI to check it), then no gate has a human holding it, and the wrong stuff sails all the way to production on green lights.** So the more it swears something's true, the more I make myself stop and verify it. Let go of this gate for a moment, and those five "safe to hand over" tasks instantly turn into landmines — because the precondition for handing them over was that this gate of mine was still standing. I got burned on this once. One time it helped me analyze a batch of data; the conclusion was pretty, the logic flowed, and because it read so smoothly I didn't dig in — I wrote it straight into a report. Only later did I find out it had mixed up the definitions of two metrics, and the whole conclusion was backwards. After that I set myself an iron rule: **the smoother the conclusion, the more I stop; the more confident it is, the less I let myself get lazy.** **Number three: the things between real people that require reading the room.** Soothing a colleague who's fuming because their requirement got cut, persuading a boss who flatly disagrees, finding balance between two teams that have started fighting — these I've never thought about handing to AI. AI can help me soften the wording of an email, but it can't read whether the person across from me is genuinely angry right now or just wants an out, whether I can crack a joke or have to stay serious. **These things — hidden in tone, in pauses, in the history between the two of you — are the hardest part of this job, and the part you least can outsource.** A polished, perfectly even-keeled email written by AI can, sometimes, do far more damage than your own clumsy but sincere two sentences. **Number four, and the most fundamental — the taste for "what is good," and the responsibility of being the one who carries it when things go wrong.** Two prototypes you can both click, and which one "feels right" and which is off — that line is something I've built up over the years; I can't spell it out but I can recognize it at a glance. And once this product runs into trouble, the person who stands up and owns it is me, not AI. **You can't make a model bear responsibility — it won't lose sleep over a wrong decision, it won't take a hit to its performance review, it won't flush red at the post-mortem.** And the PM role, in a sense, is selling exactly this: that there's a specific person, on the hook for these judgments. That's something I can't hand over, and shouldn't. ## When a new task lands, how I decide on the spot whether to hand it over The two piles above are what a year of doing this added up to, but when a task I've never done before actually lands in my hands, I don't go flipping through this list. I just ask myself three questions, and it settles fast. **First: if it got this wrong, could I catch it on the spot?** If I can catch it, lean toward handing it over; if I can't, or it only surfaces much later, keep it. If it botches a document, I know the moment I read it, so I dare hand it over; if it botches prioritization, I won't find out for months, so I don't dare. **Second: does this thing have a "standard answer" I can check against, or does it come down entirely to this moment's trade-offs, relationships, and taste?** If there's an objective right and wrong (is the material real, is the sorting accurate), hand it over; if the answer is buried in "what we actually want this quarter" or "what mood this person is in right now," keep it. **Third: if this goes sideways, does a specific person need to stand up and carry it?** If someone needs to be accountable, I don't outsource it — because responsibility simply can't be handed to a model that won't lose sleep and can't be held to account. Here's an example so you can see how to use it. A while back the question was whether to add an "AI smart reply" to a feature, and I split it into these three. Whether it works well, I'll know once I ship it and see if users click it (question one: catchable, verifiable). But "should we spend these resources on this feature, is it worth betting on this direction" — no standard answer, entirely down to our wager (question two: rides on judgment, keep), and if it really flops, I'm the one taking the fall (question three: someone's accountable, keep). So the conclusion was clear: **let AI build the prototype of that reply feature, but whether to do it, whether to place that bet — I call that myself.** Run the three questions once, and the line between hand-over and keep surfaces on its own. ## That line, it turns out, keeps sliding in one direction Lay out the two piles and you'll notice: not one of the things I kept is held there because "AI isn't smart enough yet." Quite the opposite — at writing a polished email, at laying out priorities cleanly, it may be better than I am. I kept them because their value lies precisely in being **done by a person**: a decision someone has to be responsible for, a judgment someone has to be accountable for, a relationship someone has to genuinely reach out and catch. **The scarcity of these things doesn't come from AI being unable to do them; it comes from "there must be a specific person here."** That's also why I'm not worried this line will suddenly collapse one day — it isn't drawn along "AI's capability boundary," it's drawn along "who's responsible, who has the taste," and the latter, for the near future, can still only be a human. As for the "handed over" pile, I'm well aware it's only going to grow. Today AI helps me write first drafts and dig up material; tomorrow it might get me to eighty out of a hundred on the initial solution design too. **What I need to do isn't cling to a few tasks and refuse to let go — it's keep a firm hold on that "is it right" gate, while constantly asking myself: which old task is it time to hand over now?** That said, honestly, the handing-over itself has a cost too, and this is a part I still haven't fully figured out. If AI writes all my first drafts, will I slowly forget, after long enough, how to think a PRD clear from scratch? If it digs up all the material, will I lose the ability to plunge in myself and dig out a feel for it? **There are some tasks I still insist on doing from scratch myself once in a while — not entirely because AI does them badly, but because I'm afraid that muscle will atrophy. When the moment comes that I need to judge whether it did something right, I'm afraid I'll have lost the feel for it.** Beyond hand-over and keep, there seems to be another question — how much to hand over, and how much to keep for my own practice — and that one I'm still figuring out. I'd like to hear how you weigh it, too. --- # AI-Era PM Interviews: How to Answer the 5 Questions They Love Most URL: https://doaipm.com/en/blog/pm-ai-interview-questions/ Published: 2026-07-15 Tags: product manager, interviews, job hunting, AI product manager, career, hands-on I've interviewed a lot of product managers these past two years, and been interviewed myself. One pattern is so stark I still remember it: **the moment an AI question comes up, eight out of ten people immediately start reciting concepts** — what RAG is, the difference between fine-tuning and prompting, how the attention mechanism in a Transformer works. The smoother the recital, the colder I feel inside, and by then I've basically decided I won't hire them. Not because they got it wrong. Because **these questions were never testing what you memorized; they're testing whether you can think.** The instant your mouth opens with a definition, you've told the interviewer: I studied this as a fact to memorize, but I've never actually made a decision on it. This is exactly where the AI PM role differs most from the traditional one. With traditional software, once you've thought it through, it more or less behaves that way. AI doesn't — **a model can dazzle in testing and then fall apart in production; it works great for 90% of people and talks nonsense to the other 10%; and often you can't even explain why it got something wrong, let alone patch it and move on.** So what the interviewer actually wants to know isn't whether you know the terms — it's how you make product decisions in the face of something that *will* be wrong, is hard to explain, and is hard to fix. The 5 questions below are the ones I've heard, and asked, most often these past two years. I won't hand you a standard answer — most AI questions don't have one, and the interviewer is grading your reasoning process, not your conclusion. I'll just tell you what each question is really measuring, how I'd answer it, and the kind of answer most likely to crash and burn. ## 1. "This AI feature is sometimes wrong. How do you decide whether it's ready to ship?" This is almost always the first question in an AI PM interview, and it's the dividing line. **What it tests: can you accept the premise that AI will definitely be wrong, and still make a responsible decision?** The traditional-software instinct is "there's a bug, so fix it until there are no bugs," but with an AI feature you can never drive the error rate to zero. If your answer carries even a whiff of "I want to get accuracy to 100% before I ship," you're basically out — it means you haven't stepped into the world of AI yet. How I'd answer: **I don't chase zero errors; I chase "when it's wrong, I can absorb the cost."** The first thing I ask is — if this feature is wrong, what's the worst that happens? An AI that drafts emails for users: if it's wrong, the user tweaks it, the cost is tiny — so I'll ship at 80% accuracy, because the other 20% the user can cover themselves. But an AI that auto-charges a user's card, or suggests a diagnosis to a doctor: one mistake is an incident, so I wouldn't dare ship even at 99% without adding human review, adding fallbacks, adding a "when in doubt, don't act" escape hatch. **The same accuracy number can be shippable or not — it depends on the cost of being wrong, not on the number itself.** The answer most likely to crash: fixating on the accuracy number alone and never talking about what happens when it's wrong. Treating "is the AI good?" as a test score instead of a product question of "can I survive the worst case?" — that's 2023 thinking, and a 2026 interviewer hears it and knows instantly you've never actually shipped an AI feature. ## 2. "For this requirement, do you use prompting, RAG, or fine-tuning? Why?" This one comes up absurdly often, especially for roles tied to large models — prompting / RAG / fine-tuning, pick one, is nearly guaranteed. A lot of people think it's testing technical knowledge. In fact **it's testing whether you can make tradeoffs.** The interviewer doesn't expect you to hand-write fine-tuning code. What they want to see is: given a requirement, can you rank these paths by cost, speed to results, and controllability, and then explain clearly why you chose this one and what you're on the hook for by dropping the others. How I'd answer: **I always start from the lightest option and add weight only as needed — if prompting can solve it, I will not reach for fine-tuning.** Because prompting is the fastest to change, the cheapest, and if it doesn't work today I can tune it tomorrow. RAG fits the case where "the answer has to be grounded in my own set of documents, and it has to stay updatable." Fine-tuning is the heaviest — training costs money, needs data, and each change has a long cycle — so I only move to it when "I've tried the first two, the results genuinely aren't good enough, and this capability is worth spending that cost on." The key is that last part — **you have to be able to say what each path gives up.** Choose prompting, and you give up stability (the same question right today, wrong tomorrow). Choose fine-tuning, and you give up flexibility (want to change a behavior, you retrain). If you can lay out "I picked A, the cost is losing B, but for this requirement that cost is worth it," you've aced the question — even if you can't write a single line of model code. The answer most likely to crash: opening with "well, it's got to be fine-tuning, that gives the best results." Someone who reaches for the heaviest option out of the gate — the interviewer assumes you don't understand cost, and have never sweated over a real project's budget. ## 3. "Tell me about a time you decided NOT to use AI." This one is quietly brutal, but it filters people beautifully. These past two years everyone everywhere has been shouting about AI, and a product manager who has **never once said "AI doesn't belong here"** is probably chasing a trend, not building a product. **What it tests: do you treat AI as the goal, or as a tool?** The interviewer wants to confirm you won't cram an AI feature into a place that plainly doesn't need one just because "the boss wants AI" or "it makes a better fundraising story." How I'd answer: I'd tell a specific one — say, a feature where everyone wanted to bolt on "smart recommendations," and I blocked it. Because in that scenario users had only a handful of fixed options; a hard-coded rule was faster, more accurate, and never wrong. Forcing a model on top would be slow, expensive, and occasionally skew the recommendation — pure AI-for-AI's-sake. **In the end we solved it with the dumbest possible if-else, and I think that was one of the most right decisions I made that year.** Having one concrete "I said no to AI" story is worth more than dazzling anyone with RAG. Because it proves the one thing an interviewer wants most and can test least: **you have judgment, AI doesn't order you around, you're the one using it.** The answer most likely to crash: "I can't think of a case like that — I feel like AI improves the experience pretty much everywhere." That single sentence files you under "AI believer" — and no mature team wants a product manager who can't tell when AI shouldn't be used. ## 4. "How do you trade off cost and latency?" A few years back this was treated as an engineering problem — how many tokens you spend, how slow the response is — that's for the backend to worry about. **But the bar changed in 2026: cost and latency sit right on the AI PM's own dashboard, alongside quality and experience, to be weighed explicitly when you're building the roadmap.** **What it tests: do you understand that an AI feature burns real money per call, and that being one second slower can lose you a batch of users?** Once traditional software is written, one more user costs almost nothing extra. AI isn't like that — every single call spends money, and the more it's used, the more it burns. A product manager who doesn't watch this ledger will build a feature with a great experience that the company can't afford to keep alive. How I'd answer: I treat it as an explicit three-way tradeoff — quality, cost, speed — and you usually can't max out all three. I ask what this feature actually wins on: if it wins on "answering accurately," I'll let it be slower and pricier and use a stronger model; if it wins on "grab it and go," I might pick a cheap, fast small model and sacrifice a bit of quality to buy back response speed and cost. **The point is I have to know what I'm trading for what, instead of defaulting to the strongest, most expensive option every time.** The answer most likely to crash: "Just leave that to engineering to optimize." Punting cost and latency to engineering — that's precisely the wrong answer the interviewer is waiting for, and the moment you give it, the label "this person doesn't understand the economics of AI products" is stuck on you. ## 5. "Why do you want to be an AI product manager?" It looks like a polite icebreaker; it's actually an honesty test. **What it tests: have you really put your hands on this, or were you drawn in by the hype and the salary, ready to recite a script?** Because every technical follow-up that comes right after will check whether this "why" of yours is real. How I'd answer: I won't give the "because AI is the future, it's the inevitable trend" kind of correct-but-empty line — the interviewer hears that twenty times a day, so it says nothing. I'll tell one specific small thing: the first time I used AI to build something I'd been sitting on forever and couldn't code myself, and got it real and clickable in a single afternoon — that jolt of "wait, one sentence from me and it comes true." **A real, specific, slightly clumsy first time is always more convincing than one grand, correct pronouncement.** The answer most likely to crash: reciting industry trends, reciting big words. The bigger and more correct you make it sound, the more certain the interviewer is that you've never actually done it — because people who have done it talk in specifics, not in pretty pronouncements. ## Don't forget — the real exam is in the follow-ups For every question above, finishing the first round isn't the end. Where an AI PM interview really tests you is the interviewer's next, offhand "**Then what?**" You say "I'd dare to ship this feature at 80% accuracy," and they follow up: "So how exactly do you cover the 20%?" "If it turns out to be only 70% in production, what do you do?" You say "I chose prompting over fine-tuning," and they follow up: "So if no amount of prompt tuning gets you past 60%, when do you change your mind and move to fine-tuning? Where do you draw that line?" **They're not trying to trap you; they're confirming whether that pretty conclusion of yours was thought through or memorized.** Memorized answers fall apart two follow-ups deep — because a memorized answer has no next layer; you're left holding a single, isolated conclusion and can't answer "if the situation changed, how would I adjust?" A thought-through one gives you something to say no matter how far they push, because you've actually walked that path to the end in your head. So when you prepare, don't just prepare answers — **ask yourself "then what?" one more time**: if this decision turned out wrong, how do I recover; if conditions change, when do I change my mind; what gives me the right to draw this line here. Walk the follow-ups yourself first, and when the real "then what?" lands in the interview, it stops being a hurdle and becomes your chance to show what you've actually got. ## In the end, it's all measuring the same thing Put these 5 questions side by side and you'll see what the interviewer is weighing again and again is the same thing: **do you have a concrete story in hand — a real decision, a real tradeoff, ideally with a number attached?** What accuracy you'd dare ship at, why you chose prompting over fine-tuning, the time you blocked an AI feature, whether you traded quality for cost or cost for speed, what made you jolt the first time AI stunned you — on all five, the people who answer well are all telling you about things they actually did, and the people who answer badly are all reciting definitions and trends. **Hollow stories lose the offer; concrete stories win the offer** — and in the AI PM role, that line bites harder than in almost any other. So if you're preparing for this kind of interview, my advice isn't to memorize a question bank. Go back and dig out the AI-related things you actually did, one by one, and think each one through clearly: what the decision was, what you gave up, how it turned out, whether there was a number. Line up five or six stories like that, and no matter how these questions get reworded, you'll always have something real to say. Here's one thing I'm still turning over, and I'll toss it your way: when AI drops the barrier to "make a thing" this low, the "what have you built?" questions in interviews will get easier and easier to answer — because everyone can build something now. So what becomes the question that actually separates people then? My guess is it drifts toward "what did you **not** build, and **why not**." But that's only a guess. I haven't seen the answer yet either. --- # A Day as an AI-Era PM: How I Turned One Sentence Into a Prototype You Can Actually Tap URL: https://doaipm.com/en/blog/pm-one-sentence-to-prototype/ Published: 2026-07-14 Tags: product manager, AI workflow, prototype, career skills, high-fidelity, hands-on Let me start with something concrete. Last Wednesday afternoon, a fuzzy idea popped into my head: "I want a little thing to jot down what I spend." A little over two hours later, my coworker was tapping away on my phone for real — tap "Add," type an amount, pick a category, go back to the home screen and see how much he'd spent this month, with a pie chart split into a few slices. Not a single line of code. I'm not telling this story to show off how magical AI is. I'm telling it to make a counterintuitive point: **most people think that now AI is here, PMs had better hurry up and learn programming or they'll get replaced. But my gut sense over the past six months or so has been exactly the opposite — the barrier to writing code is collapsing fast, and the thing that's actually becoming valuable, and scarce, is something else entirely: getting a thing described clearly enough that AI gets it right on the first try.** I'm not going to argue that point in this piece — I'm going to show you exactly how I did it. I'll use that little expense tracker as the example and walk you through it step by step, potholes included. If you've got an idea you've been sitting on for ages, by the end you should be able to go try it yourself. ## Step one: don't rush to make it build — make it interrogate me first When I first started building things with AI, my favorite mistake was to fling that fuzzy sentence in my head straight at it — "build me an expense-tracking app" — and then wait for a finished product. The result was always the same: it would hand me something *it* thought I wanted, a thousand miles from what was in my head. I'd correct it a little, it'd drift a little, and the two of us would sit there guessing at each other until I lost my temper. Later I changed one habit — it's just one sentence, but it works absurdly well: **I stopped asking it to build directly, and started making it interrogate me first.** I'll say: "I want to build a little expense-tracking tool. Don't write anything yet. First ask me 5 questions — the things you need to know but that I haven't told you yet." So it asks: who's it for, just you or multiple people? Do you want categories, defined by you or a few presets I pick? Should amounts separate income from expenses? Is storing data locally on the phone fine, or do you need to see it across devices? Do you want a budget alert? **See, every one of these questions lands exactly on the stuff that was a muddled mess in my head — stuff I hadn't thought through.** It forced "what do I actually want" out of me. By the time I've answered those 5 questions, the shape of the thing is clearer in my own mind. And *now* when I let it build, the odds of getting it right in one shot go way up. A pothole I stepped in: **the two minutes you save by skipping this step, you'll pay back with two hours later.** Once, I got annoyed at all its questions and just told it to build. It produced something fully featured and completely not what I wanted, and by the end the rework was worse than starting over. Now I never skip this step. ## Step two: change only one thing at a time The first version comes out, you can tap it, but it's definitely off in some way. And here's where the second pothole is waiting: **dumping all ten things you don't like on it in one breath.** "The home screen color's too pale, the category icons are ugly, I want to change the pie chart colors, the Add button's too small, oh and can the amount auto-add the decimal point, and while you're at it throw in a search…" I used to do exactly this, chasing speed. The result: it makes the changes, gets the color right, but forgets the button; or fixes the button and breaks the pie chart. The more changes at once, the more it drops one to catch another, and you can't even tell which of your sentences did what or which one it quietly skipped. **Now I've switched to saying one change at a time — say it, tap it once on the phone, confirm that one spot is right, then say the next.** Slow? Looks slow. But every step is solid, no backtracking. I've done the math: going through it this way is actually much faster than "say everything at once, then rework it all together," and I know the exact state of things the whole way through. This isn't some AI trick. It's the plainest rule in product work: **small steps, each one verifiable.** It's just that AI has crushed the cost of each step so low that you have even less excuse to get greedy. ## An aside: how to phrase each change so AI actually gets it I said "change one thing at a time" above, but just "one" isn't enough — the key is *how* you say that one thing. This is the most hands-on part of the whole piece, and the one that pays off fastest. The biggest pothole I stepped in was **giving AI instructions with adjectives.** "Make the home screen nicer," "make the button bolder," "make the color classier" — the moment those words leave my mouth, I've handed all the judgment back to it, because "nice" and "classy" have ten thousand interpretations in its head, and it'll grab one at random, most likely not the one you meant. So I forced myself to do two things instead: **give it a reference, give it a state, don't give it adjectives.** Give it a reference — instead of "make it nicer," say "the cards on the home screen, use whitespace like WeChat Read does, three or four per screen, don't cram them." It knows immediately what you want, because you've given it a concrete thing to align to instead of a vague verdict it's free to interpret. Give it a state — instead of "handle the case where there are no entries," say "when there isn't a single entry, show one line of gray text in the middle of the home screen: 'No entries yet — tap the button at the bottom right to add one,' and don't show an empty pie chart." **Spell out exactly what the screen should look like in each state** — what it looks like with data, without data, while loading, when it errors. The more specific you are, the less room it has to "freely interpret" its way into something you don't want. My self-check now is: **after I finish an instruction, I look back for adjectives.** If there's a "nice," "bold," "classy," or "clean this up," I stop and translate it into "reference what, look like exactly what." This little move has done more for me than any AI trick I've learned. ## Step three: the first version runs on real data — no "placeholder text" This is the one I think gets ignored most and affects the outcome most. A lot of people building prototypes have AI put up a "frame" first — filled with fake stuff like "Title 1," "content placeholder," "¥000.00" — thinking "get the structure right, fill in content later." I don't do that anymore. **I have it run on real data from the very first version.** For this expense example, I just had it preload a few things I actually spent yesterday: breakfast 12, a cab 28, groceries 63, and one that stung a bit — 2000 for my kid's classes. Why? Because **fake data lies to you.** When everything is "¥000.00," the interface looks clean and tidy and you think "yeah, this is fine." But the instant you drop in a real, uneven-digit number like "2000," the problems all surface at once: the amount gets long and the right edge butts up against the screen; the pie chart gets crushed by that one 2000 entry until the smaller slices are nearly invisible; a slightly longer category name warps the whole layout. **Every one of these potholes is invisible with fake data, and they all smash into real users' faces the moment you ship.** With real data, they're exposed in your own hands in the very first version — and the earlier they're exposed, the cheaper they are to fix. At this stage it's a one-line change; fixing it after users are cursing at the live app is a whole different thing. My habit now: if I can use real numbers I'll never use a placeholder, and the more real and more extreme the better — the longest name, the biggest amount, the emptiest state (what does the interface look like when there isn't a single entry?) all need to show up in the first version. ## Step four: you must tap through it yourself on a real device Once the prototype's built, AI will usually tell you very confidently, "Done, all the features are implemented." Don't believe a word of it. **It's not lying — it just doesn't have hands. It can't actually tap anything.** I've been burned by this. Once it swore up and down that recording, deleting, and stats were all done, I couldn't find a flaw in the code logic either, so I believed it. Then my coworker took it, added an entry — fine; tapped that entry to delete it — nothing happened. It had drawn the delete button, but the "actually delete it when tapped" action, it had missed — and it had no idea, and still told me "done." Ever since, I set an iron rule: **any "it's done" only counts after I've personally walked the key path through it on a real device.** Add an entry, watch it appear in the list, delete it, watch the stats change with it — until I've tapped that whole chain through myself, I treat it as unfinished. This step takes under three minutes, but it's the only wall standing between you and "discovered by users after launch." I've got a dumb little method for how to walk it, too: **before I start, I write down the two or three most critical paths of this thing on a piece of paper.** For this expense tracker, I wrote: "① can add an entry and see it ② can delete an entry ③ home-screen numbers and pie chart change accordingly." Just those three. When it's done, I don't look at what AI says — I take that paper and tap through it on the phone, one line at a time. Cross off each path that works; whatever I can't cross off isn't finished. Don't underestimate that piece of paper. It forces you, before you start, to get clear on "what few things does this thing actually stand or fall on" — a lot of people lose the plot halfway through precisely because they never pinned those two or three main paths down on paper. They build and build, get dragged off by details, and end up with a pile of features while the single most core path doesn't even work. **Get clear on which few paths must work, then let AI build, then verify those paths by hand** — those two pieces of paper, the front end and the back end, matter more than however much code it wrote in between. ## Step five: the acceptance line is "can you tap through it," not "does it look right" String the steps above together and it's really the same judgment showing up over and over: **what, exactly, do I use to decide whether this thing works?** The answer I've given myself: **whether you can actually tap through one complete path, not whether it looks right.** "Looks right" is the easiest thing to be fooled by. Drop a screenshot in the group chat and everyone says it looks great — and maybe not one person has actually recorded a single entry. "Can tap through" is hard: a real person, from opening it, to recording an entry, to seeing the result, gets through the whole thing without getting stuck, without a single dead button — once that path works, the thing genuinely stands. This is also where I think AI-era PMs should hold the line hardest. **When the cost of building is crushed to nearly zero, "producing something that looks presentable" is no longer worth anything — the screen is full of them. What's worth money is whether you can still judge: is this presentable-looking thing actually usable, or does it just look like it?** That judgment, AI can't do for you — because it's the very thing that will confidently say "done" while missing the delete button. ## So — do you actually need to learn to code? Back to the question at the top. My answer is: **you don't need to learn how to write code, but you do need to learn how to get your words clear enough that AI gets it right on the first try — and those are two different things.** The former is learning a craft that's depreciating; the latter is training a judgment that's getting scarcer: forcing a fuzzy idea into clarity, advancing only one step at a time, checking it against real things, tapping it through by hand before you believe it. None of these, honestly, is "technical." They're more like the habits a person should already have — someone who's clear on what they want and willing to verify it step by step. AI just amplifies the payoff of those habits many times over: the clearer you think, the more accurate what it gives you; the more you fudge it, the more it fudges you back. There's still stuff I haven't figured out, and I'll toss it your way: when "turning an idea into something you can tap" gets fast enough to run several rounds in a single afternoon, where exactly is the line between a product manager and "an ordinary person who's clear on what they want"? I'm still looking for the answer myself. But at least I know the line isn't drawn at "can you write code" — that little expense tracker last week, from start to finish, I didn't touch a single line of code. --- # Kung Fu Women's Soccer only scored 6.6, yet Stephen Chow is the most ruthless product manager I've ever seen URL: https://doaipm.com/en/blog/stephen-chow-scored-the-wrong-product/ Published: 2026-07-13 Tags: product judgment, user needs, Stephen Chow, Kung Fu Women's Soccer, commercialization, tech commentary Let me open with a question I can't answer cleanly myself: **if you had a product that your professional users rated 6.6, that 8.6% of them slapped with one star, that got roasted from top to bottom in the comments — but that sold 500 million in two days, should you panic, or should you quietly celebrate?** This isn't hypothetical. Kung Fu Women's Soccer, which opened in theaters over the last couple of days, is exactly that thing. Directed and written by Stephen Chow, starring Zhang Xiaofei, Dilraba, and Zhang Yixing. It opened at 6.6 on Douban, with nearly 80,000 ratings and 8.6% giving it one star. The complaints are frighteningly consistent: effects that look like they were smeared together by AI, mainland actors overacting their impression of Chow's absurdist comedy, and a plot that's just Shaolin Soccer with the genders swapped — "reheated leftovers" filled up the whole screen. But the numbers on the other side look like this: it crossed 100 million in 27 minutes on day one, took a 76.8% share of screenings, and broke 500 million at the box office in two days, with total box office forecasts revised up from the original 1.428 billion all the way to 1.865 billion RMB. In all my years doing product, the thing I dread most is exactly this kind of split between rating and revenue. Because it forces you to admit something uncomfortable: **the "make a good product" standard you believe in may simply not be the one the market is paying out on.** The 6.6 is one signal, the 1.8 billion is another, and they point in opposite directions. Which do you trust? My answer is: the two numbers don't actually contradict each other, because **the critics and the box office are scoring two different products.** ## The critics are grading the craft; the audience is paying for the feeling That 6.6 on Douban is a grade for "the film" as a piece of craftsmanship: are the effects sharp, is the acting relaxed, is the story fresh, are the shots considered? That standard isn't wrong — it's the industry's quality-control line for the craft of filmmaking. Measured against that line, Kung Fu Women's Soccer does come up short at every turn: rough effects, stiff acting, a stale story. None of it is unfair. But of the 80,000, 800,000, or 8 million people who paid to walk into a theater, the overwhelming majority weren't there for "a good movie." They were buying something else: **Stephen Chow.** More precisely, "that Shaolin Soccer feeling" — the summer of 2001, the line "if a person has no dreams, how are they different from a salted fish," a ticket back to a youth they can never actually return to. The first time I understood how deadly this is was many years ago, working on a utility app. Our team spent the better part of a year rebuilding the interaction, convinced it was far cleaner, more modern, and more professional than the old version. Launch day, everyone was waiting for the praise. Instead a wave of longtime users showed up to yell at us: where's that ugly, cramped old version? Where did you hide the feature I use? I spent an entire quarter chewing on that lesson — **I had been grading the product with "the good as I define it," while users were paying based on "is the thing I want still here." Two different rulers, and I'd grabbed the wrong one.** Stephen Chow didn't grab the wrong one. He knows all too well that what his users want isn't "new" — it's "that feeling." So he's willing to let the plot be a gender-swapped remake of Shaolin Soccer, willing to let the jokes still be the same absurdist bits from twenty years ago — **on the critics' ruler this is called "no innovation"; on his ruler, it's called "delivering with precision the thing users actually came to buy."** ## The people rating it and the people buying tickets are two different crowds There's another layer, one that stings more than the "two rulers" point: **the people carefully rating it on Douban and the people handing over money at the box office overlap far less than you'd think.** The kind of person who goes to Douban to write a long review, gives a commercial comedy one star, and picks apart the effects frame by frame — that's usually a small slice with a strong urge to express themselves and real aesthetic demands. They're loud, they write sharp, they spread wide, and so you get the illusion that the whole world is trashing this movie. But the people who actually determine that 1.8 billion are a different, much larger crowd — they don't write reviews, they don't rate, some don't even have a Douban account. One weekend they take their parents or their kids out for a laugh and buy a ticket. **The former manufactures the reputation, the latter manufactures the box office, and these two groups are often not the same group.** I got burned on this too. That redesigned app I mentioned, the one that got roasted — at the time the forums and the app store reviews were a sea of complaints, and a few of us stayed up night after night staring at those bad reviews, fixing them one by one, getting more anxious the more we fixed, convinced the product was done for. Then we pulled the data and found: the ones complaining loudest were a few hundred high-frequency longtime users; and in a place they couldn't see, hundreds of thousands of new users who'd never made a sound were quietly using the new version, with retention that was doing just fine. **We nearly changed away the thing that hundreds of thousands of silent people wanted, just to soothe a few hundred of the loudest voices.** From then on I learned one thing: a product's loudest voice and its biggest wallet are often not in the same place, and you have to be clear about who you're making the decision for. Stephen Chow clearly is clear about it. He didn't try to shoot the film "highbrow" to please that 8.6% — he knows that 8.6% was never the group he needed to win this round. He was aiming at the silent majority — the ordinary viewers who don't go to Douban and only recognize the "Stephen Chow" brand. **To keep your nerve under a screen full of bad reviews and still place your big bet steadily on the silent users — that read on the user base is far harder to come by than knowing how to make a movie.** ## "Reheated leftovers," in the product world, is actually a compliment I know "reheated leftovers" is meant as an insult. But translate it into product language and it immediately changes flavor — it's called **reusing a repeatedly validated product framework.** The top products are all "reheating leftovers." Every year the iPhone keynote gets trashed as "they just swapped the camera," yet on that very "unoriginal," steady iteration it walks away with the vast majority of the industry's profit. WeChat is a decade old and its core is still those same few functions; the thing Zhang Xiaolong is proudest of is precisely "what we didn't build." Coca-Cola has sold one formula for over a hundred years. **What these things have in common: they found a framework that genuinely works, and then had the discipline not to touch it.** "Reheating leftovers" and "reusing a framework" are literally the same act; the only difference comes down to one question — **is the framework you're reusing still working, or not?** Is the Shaolin Soccer framework still working? Twenty-five years on, its name can still get several million people to willingly pay to walk in, and its day-one screening share could be pushed up to 76.8% — the distributor dared to stake screenings that heavily precisely because this framework has been validated by the market over and over and almost never comes up empty. From a pure business-betting standpoint, reusing a framework that hasn't failed in 25 years is far lower risk than gambling on a brand-new, unvalidated idea. It's not that Stephen Chow can't make something new — he **worked out the odds on this bet**: nostalgia is the highest-win-rate card in his hand. Yesterday I wrote a piece about riding out a typhoon, and there's a line in it I want to say again today: a real master doesn't bet on the weather, he bets on the constants. Novelty is weather — it turns on a dime; the meme that lands this year is dated by next. But nostalgia, sentiment, "wanting to see that feeling from my youth one more time" — those are constants of human nature that don't change for decades. **Stephen Chow's bet isn't on "will the audience like something new," it's on "do people get nostalgic" — and he knows the answer to that one with his eyes closed.** ## He also got right the one thing PMs are most likely to overlook: the channel Having the right product isn't enough. The real kill shot in the Kung Fu Women's Soccer campaign is that 76.8%. What does a 76.8% day-one screening share mean? It means that on that day, if you walked into any theater, it was hard to watch anything else — nearly everything in sight was this film. Crossing 100 million in 27 minutes wasn't just about the content; it was about getting this product **right in front of every user's eyes while sealing off most of the other options.** And he stacked one more layer on top of this — **timing.** The film was slotted into the summer season, the window with the thickest moviegoing traffic and the most families going out all year. The same product, dropped in a cold window versus staked on the summer season, can differ by several times in volume. Choosing the right launch window and then using screening share to eat that window whole are two moves stacked together: **on the day with the biggest foot traffic, standing at the intersection with the biggest foot traffic, stacking your goods in the most visible spot.** This is the link PMs are most likely to underrate. We love to spend 90% of our effort polishing the product itself, figuring "if the thing is good, people will come naturally," and then we phone in the "dirty work" of distribution, channels, shelf placement, and launch timing — and end up with a good thing rotting in the warehouse. Stephen Chow's team did the reverse: the product (nostalgia) is the safe, validated card, but they took the channel (screening share) and the timing (summer season) to the extreme. **A validated product, paired with a channel that fills every shelf — that's the real recipe for those 500 million in two days.** Content is only half of it; the other half is that 76.8% a lot of people are too proud to calculate. ## But I don't want to write this as a hymn to Stephen Chow — he's overdrawing something If I wrapped up here, this would turn into a feel-good rant of "ratings don't matter, making money is the real skill." But after doing product long enough, the first thing I learned is: **behind any beautiful number, you have to ask, "where was this money moved from?"** A big chunk of this 1.8 billion for Kung Fu Women's Soccer wasn't earned by the film itself — it was earned on the trust that the name "Stephen Chow" has accumulated over more than twenty years. That trust is his real product, his moat, the brand asset no one can copy. And every time he "reheats leftovers," he's withdrawing money from that account. That 8.6% of one-star ratings, I don't see as ordinary bad reviews. I see it as **the sound of the brand asset starting to pay down its debt early.** This time, the longtime audience was still willing to pay for nostalgia, but some of them, walking out of the theater with that line "this time it really was a bit of a cop-out," are quietly docking points from the next move. The cruelty of nostalgia is this: it can be cashed out for very high box office all at once, but it's a consumable — every withdrawal leaves less, and you can't withdraw like crazy and still expect the balance not to move. So in my eyes, Stephen Chow isn't someone who "got lucky and made money" in a fog. Quite the opposite — he's a top product manager who **soberly made a trade with a price attached.** He probably knows better than anyone that the craft this time doesn't deserve this box office, and knows that every reheat thins the brand another layer. He just, after running the math, chose to trade a portion of long-term brand asset for this one certain, enormous slug of short-term revenue. That's a cool-headed trade-off, not a fumble. **Someone who can work out user needs, framework reuse, channel, and brand overdraft all at once — and still dares to place a big bet — I find it hard not to call a top product manager**, even though I don't actually like the thing he handed over this time. ## That ruler — you've got one in your hand too There's one thing I still haven't fully figured out, so I'll just toss it to you: how many more times can the nostalgia card be played? Whether Stephen Chow's balance can be calculated by him, or whether it takes one film really flopping to find the bottom — I don't know. But there's a question I think every person doing product should ask themselves: **that thing of yours that people dismiss as "unoriginal, same old stuff" — is it truly ready to be retired, or have you finally found the framework you don't need to change, and you're just being held hostage by "I must innovate"?** From the outside these two look identical; the difference is in one spot — **is the framework you're reusing, dropped into today's market, still working or not?** Working, and it's Coca-Cola; not working, and that's when it's reheated leftovers. Stephen Chow bet this time that it's still working, and the box office answered "yes" for him. As for next time, that 8.6% of one-star ratings is already keeping the tab for him. --- # A Typhoon That Fizzled Out: How a PM Survives the Darkest Hour Like Riding Out a Storm URL: https://doaipm.com/en/blog/typhoon-darkest-hour/ Published: 2026-07-12 Tags: darkest hour, crisis response, product manager, postmortem, typhoon, tech commentary Let me start with something that might rub people the wrong way: **the vast majority of "darkest hours" you'll live through will end exactly like this typhoon — a false alarm. And those three words, "false alarm," can kill a product manager faster than the typhoon itself.** This year's Typhoon No. 9, "Bawei," veered south last night. Instead of slamming straight into Zhoushan, it weakened to typhoon grade and headed for the coast from Wenling in Zhejiang down to Xiapu in Fujian. So here in Zhoushan and Putuo, the ferries that shut down on July 9, the cancelled flights, the boatloads of fishing vessels moved overnight from their anchorage to the west pier at Shenjiamen — in hindsight, it all looks like wasted effort. This morning someone in the group chat was already saying: "If I'd known, I wouldn't have bothered sailing the boat out in the dead of night." Stop on that sentence. Because the part of this whole thing most worth chewing on is hidden right inside that "if I'd known, I wouldn't have bothered." In the last piece I wrote about Putuoshan, my point was: don't go pray to Guanyin to block the typhoon for you — go fix the seawall. **Don't credit "not getting hit" to some mysterious protection; that's a misattribution.** This piece takes one step further: the seawall is built, the alarm has actually sounded — so on that night, what exactly should you, the person running the show, do? My answer: riding out a storm was never one move; it's three moves in a row — **before the typhoon arrives, prepare fully; in the thick of it, take the wind and rain; after it leaves, clean up the trash.** These three things — even if nine times out of ten it turns out to be a false alarm — you can't skip a single one. That's how a product manager survives the darkest hour. ## 1. Before the typhoon arrives: your composure was all stockpiled long ago The most deceptive thing in a darkest hour is "composure." You'll see people whose systems crash in the middle of the night, whose data goes wrong, whose boss is calling — and they can still sit still and untangle the mess one thread at a time. You assume it's innate temperament, improvised on the spot. It isn't. **Not one gram of composure in a darkest hour grows on the spot; every bit of it was stockpiled, little by little, on all those "days without a typhoon."** For someone who builds products and systems, what you "stockpile" is very concrete: a rollback switch you can hit at any moment; a canary rollout that lets you push 1% of traffic out to test first, instead of shoving 100% of your users into the water all at once; a ring of monitoring and alerts so you know where you're bleeding before your users start cursing; a contingency plan that's written down and **actually rehearsed**; and the trust and cash-flow buffer you built up bit by bit while you didn't yet need it. The key phrase here is "actually rehearsed." On January 31, 2017, GitLab had an incident that people still bring up to this day. Late at night, an engineer ran a delete command on the wrong database and wiped the production data. Dropping the database wasn't the end of the world in itself — what really sent chills down everyone's spine was what came next: they had five backup and replication mechanisms, and when they went down the list, **not a single one actually worked.** The automated pg_dump backup had never actually run successfully because of a configuration error; a failed backup was supposed to trigger an email alert, but those alert emails were silently rejected by the DMARC settings, so no one ever got them; there were Azure disk snapshots, but restoring from them would take over 18 hours. What finally pulled GitLab back from the edge of the cliff was a snapshot one engineer had **happened** to take by hand six hours before the incident. It was on that one "happened to" that they lost only six hours of data. This incident laid out, in bloody detail, what "prepare fully" really means: **you thought you'd built a seawall, and only when the typhoon came did you find out it was painted on.** Five backups — sounds impregnable, yet not one had been seriously verified, seriously rehearsed. An un-rehearsed contingency plan isn't a plan; it's a wish. That line you wrote in the doc — "we have a comprehensive disaster-recovery solution" — is, before the typhoon comes, no different from a blessing. An even harder-nosed approach is to not wait for the typhoon at all. Netflix keeps a program called Chaos Monkey, and its daily job is **to reach into the production environment, pick a calm sunny afternoon, and randomly kill a few machines that are actively serving traffic.** It sounds like self-harm, but it's actually the most honest kind of rehearsal — rather than wait for a real typhoon to reveal that the seawall is painted on, you turn a mad monkey loose in your server room every single day, crash a not-yet-broken system on purpose, and see whether it can actually patch itself back up. The composure that survives Chaos Monkey is real composure; the composure that's never been put through this is just luck that hasn't run out yet. So the first layer of "false alarm" shows up right here: those rollbacks, rehearsals, and backups you stockpiled — the vast majority of them, you'll never use even once in your life. Never using them makes you feel it was all for nothing. But flip it around: **those peaceful days you had weren't good fortune — they were bought with preparation.** Someone who "never wastes effort on preparation" simply hasn't had his turn yet. ## 2. In the thick of the typhoon: you have to stand out front in the rain No matter how thorough your preparation, the moment the typhoon actually hits you in the face, it still hurts. What's being tested now is something else: taking ownership. In the thick of a crisis, there are two things the person running the show absolutely must not do — **run, and pass the blame.** Running is hiding yourself away and waiting for it to blow over; passing the blame is rushing to find "who pressed that button." Both moves make you feel a little better in the moment, but they simultaneously kill the most expensive thing there is: whether your team still dares to charge forward. Taking ownership is doing the opposite — **stepping out to the very front and making the decision that's hardest right now but has to be made immediately.** Stop the bleeding first: roll back if you should roll back, take it offline if you should take it offline, cut it if you should cut it, admit the mistake if you should admit it, notify users instead of dragging it to tomorrow. The tab — you own it first. Back to GitLab that day. They did something that's still counterintuitive even now: **they ran the database recovery on a public livestream.** Thousands of people watched online as they laid out the embarrassment of the deleted database, the fact that all five layers of backup were useless, and fixed it back piece by piece. This wasn't a show. This is the hardest form of "taking ownership" — I screwed up, I'll fix it in front of everyone, I won't hide. When a company dares to do this, the engineers on the team come to know: when something breaks here, the first reaction is to fix it, not to hide it. Contrast that with the kind of place that calls a blame meeting the moment something breaks. An incident hits, and the first thing everyone does is figure out how to clear themselves — quick, back up the logs, screenshot the chat history, "that module wasn't mine" out of the mouth. **The opposite of taking ownership isn't panic; it's deflection.** Someone who pushes responsibility down onto the team and out onto "bad luck" — after the storm passes, all anyone remembers is the sight of his back as he shrank away. This part is the hardest of the three, because the rain really is cold. Three in the morning, the situation still a mess, every pair of eyes fixed on you — that's not a good feeling. But what your team is watching in that moment was never whether you panic, whether you have the standard answer — **what they're watching is whether, when the fire is at its worst, you took a step forward or a step back.** What the darkest hour is really weighing is the direction of that step. ## 3. After the typhoon leaves: the mess on the ground is the real exam Once the typhoon passes and the sky clears, the most dangerous slackening arrives. Plenty of people weather the wind and rain, only to die in the end on "can't be bothered to clean up." The wreckage strewn across the ground after the storm is where you really show your worth. Cleaning up isn't just tidying the scene and getting the system running again — that's only recovery. Real cleaning up is **turning this round of pain into a rule you won't step on again next time**: the postmortem has to reach down to the mechanism, turning this lesson into a check that auto-triggers next time, a hidden hazard deleted, a contingency plan updated. Knight Capital died precisely on not cleaning up. On the morning of August 1, 2012, this high-frequency-trading firm deployed a new chunk of code, with an engineer manually pushing it to 8 servers — and **missing 1.** And on that very server there happened to be an old chunk of long-abandoned functionality no one had ever deleted, codenamed Power Peg. The new code reused a flag with the same name, so on that one un-updated machine, this dead code that should have been lying in its grave woke up and started frantically firing orders into the market. In 45 minutes it traded nearly 397 million shares across 154 stocks — **a pre-tax loss of $440 million.** The company was acquired by the end of that year and was gone. Look how short that chain is: **a chunk of dead code that should have been deleted but no one bothered to, plus one deployment with no verification that missed a single machine, took out an entire company.** That's the extreme of "not cleaning up" — trash left in the system doesn't vanish on its own; it just sits there quietly, waiting for the spark that lights it. Now look back at GitLab. Same kind of potentially fatal incident, and afterward they wrote a postmortem so public it was almost self-flagellating, laying out how **every single one** of the five backup layers failed, one line at a time, then fixing each one. One turned its trash into immunity; the other left the trash in place and waited for it to blow up. What each company looks like today, you already know. Cleaning up also has one spot that's especially easy to get backwards: **sweep the mechanism, not the person.** The point of a postmortem is to make "this pit gets filled automatically next time," not to dig out the poor soul who pressed the wrong button. The moment you start digging out people, the first thing everyone learns is to hide next time something breaks; only when you fix the mechanism and never ask who did it will people dare to shout the problem out the instant it happens. This "blame the issue, not the person" postmortem wasn't first invented by the internet — it grew out of aviation: when a plane goes down, the investigation board's entire purpose is to keep planes of the same model from falling out of the sky again, not to nail the pilot to the pillar of shame — because the moment you start nailing people, the next crew to make a mistake will choose to conceal it, and concealment makes the next plane crash worse. Google later wrote this into its SRE handbook and gave it a name: the blameless postmortem. The GitLab engineer who dropped the database wasn't publicly executed, and that wasn't softheartedness — it was clarity. **Hurting once isn't growth; turning the hurt into a rule you won't break again is.** ## 4. Back to "if I'd known, I wouldn't have moved the boats overnight" Now back to the opening, back to this morning's "if I'd known, I wouldn't have bothered sailing the boat out in the dead of night." That sentence is actually a mirror image of the misattribution in the last piece. Last time, people credited "safety" to Guanyin's protection; this time, people credit "safety" to "nothing was going to happen anyway." **One records the credit to the goddess, the other records it to luck — and both are quietly cancelling out your preparation for next time.** But the truth is: **the vast majority of "false alarms" are exactly what successful preparation looks like.** If an organization has never once been through a "wasted effort" in all these years, there are only two possibilities — either it's lying to itself, reading every lucky escape as "our system is fine"; or it has already been genuinely capsized once and paid the tuition in blood. Knight Capital never had a "false alarm," because it read every incident-free deploy before go-live as "the process is flawless," reading it right up until those 45 minutes. Typhoon "Bawei" turned south this time, and Zhoushan will most likely get through fine. But what lets those fishermen sleep tonight isn't the typhoon's last-minute change of heart — it's the boats that sailed out of the anchorage overnight. Whether the next darkest hour slams down head-on isn't yours to decide. All you can decide is those three things the people who ride out storms never leave out: before the typhoon, build the seawall so it can really block the waves, not paint it on paper; in the thick of it, stand at the very front of the line and take that rain; after it leaves, sweep the mess off the ground and sweep it into a rule you won't break again. Do all three, and people will say you got lucky; slack off on all three, and Knight Capital's 45 minutes will, sooner or later, come around to you. --- # Why People Insist That Typhoons Steer Clear of Putuoshan URL: https://doaipm.com/en/blog/typhoon-putuoshan-survivorship-bias/ Published: 2026-07-11 Tags: Survivorship Bias, Product Judgment, Cognitive Bias, Tech Commentary, Typhoon There's a claim that has circulated along the Zhejiang-Jiangsu coast for a very long time: **Putuoshan is watched over by Guanyin, so when a typhoon reaches it, the storm always detours around.** I'm not making this up. The internet is full of articles earnestly debating "why typhoons detour around Putuoshan," and they tell it as something miraculous — the sacred ground of the Bodhisattva of Compassion, one of China's four holy Buddhist mountains, the South Sea Guanyin statue standing right there; no matter how fierce the typhoon, once it reaches the island's edge it turns aside on its own. Plenty of people believe it, and they say it with a certain "you may not believe me, but this is simply the fact" conviction. Yet today, that claim is getting slapped down, hard. Because of this year's Typhoon No. 9, Bavi, Putuoshan's passenger ferries have been suspended since July 9, Putuoshan International Airport canceled 14 flights on July 10, and every fishing boat and pleasure craft in the Putuo district was moved overnight to the western pier at Shenjiamen to ride out the storm. Guanyin didn't make the typhoon detour around Putuoshan — she made the whole island close early. And this is with Bavi drifting toward southern Zhejiang, not even scoring a direct hit. The 2021 case needs no softening at all: Typhoon In-Fa made landfall along the coast of Zhoushan's Putuo district, and the **Putuoshan-Zhujiajian area** was ground zero — more than 6,000 meters of road were swamped by seawater. (For the record: Putuoshan and Putuo district are both part of the city of Zhoushan — they're not some other place beyond Zhoushan, they *are* the face Zhoushan turns toward the East China Sea.) So the facts are clear: **Putuoshan is no typhoon-proof zone. It stands directly in the typhoon's path, and it's taking a beating today.** Which means the question worth asking isn't "why can Guanyin block the typhoon," but rather — **why, in a place that has to suspend ferries and shut its gates for a typhoon every single year, do people still believe it's under divine protection and that typhoons detour around it?** The answer to that question is far more interesting than the typhoon itself. Because if you swap out the word "Putuoshan" and drop in "a certain great product manager," you don't have to change a single word. ## The people who get to say "the typhoon detoured" were themselves selected Start with the part that stings the most. "See, the typhoon detoured around Putuoshan again" — that sentence carries a hidden qualification to be spoken at all: **only someone from a year nothing went wrong, only a mouth that didn't happen to get blown over, gets the chance to say it.** The year In-Fa flooded Putuoshan-Zhujiajian, the people on the island were busy evacuating, patching things up, hauling waterlogged belongings to higher ground — nobody was in the mood to muse about "the Bodhisattva's protection, the typhoon detoured." And this year Bavi went south; Putuoshan suspended its ferries but wasn't flipped over head-on, so in a few days the pilgrims will return to the island and that line — "see, Guanyin watches over this place" — will surface again as naturally as ever. **Every "the typhoon steers clear of Putuoshan" you hear comes from a year that happened not to take a direct hit, and a mouth that happened not to get blown over.** The years something went wrong, the people who couldn't have said that line at the time — they don't fail to exist; it's just that their voices never make it into the sample the "legend" is built from. Statistics has a name for this: **survivorship bias**. Its most classic shape is exactly this: **the person doing the talking is themselves the product of a filter.** You think you're observing a law of nature; in fact you're only listening to the survivors — and a survivor will only ever tell you "how I made it through," never what happened to the ones who didn't, because the ones who didn't make it don't talk. Worse, memory itself is a biased filter. Over twenty years a typhoon grazes past a dozen times and occasionally comes head-on; your brain automatically files the "grazed past" ones under "see, the Bodhisattva protected us again," and files the one or two "head-on, roads flooded" times under "that year, fate was sealed" — one counted as a rule, the other as a fluke, purely on whether it flatters the thing you already wanted to believe. ## What occasionally turns a typhoon away is the subtropical high, not Guanyin Biased memory is only a passive error. **Actively crediting "not getting hit" to Guanyin is one big step further into being wrong — that step is called misattribution.** Science does have an explanation for "Putuoshan occasionally dodges a typhoon." Where a Western Pacific typhoon goes is mainly steered by the guiding airflow of the subtropical high. When the direction of that guiding airflow along the western edge of the high shifts, the typhoon turns somewhere over the East China Sea — this is a large-scale meteorological mechanism governing the entire sea area, not any single island. When Putuoshan occasionally "dodges" one, whether neighboring Taohua Island and Zhujiajian dodged it too is decided by the same guiding airflow — and it has nothing whatsoever to do with whether there's a Bodhisattva on the island or how thick the incense smoke rises. **But the human mind has a stubborn flaw: faced with a phenomenon that can be explained by statistics and physics, we would much rather believe a version that has a protagonist.** "A shift in the guiding airflow along the western edge of the subtropical high" is too cold, too devoid of warmth; "Guanyin reached out and pushed the typhoon away" has a protagonist, has will, has a story — and throws in a bonus dose of "I live in a place that's protected" security to boot. So the correct-but-boring explanation gets tossed, and the moving-but-wrong one gets enshrined. Remember this move: **taking a structural, statistical phenomenon and attributing it to the mysterious power of a special agent.** That's the entire secret of the Putuoshan legend. And this move — you can watch it happen somewhere else every single day. ## Swap "Putuoshan" for "a great product manager" The product managers we worship for "seeing the future" were run through the same filter and then enshrined by the same misattribution. Steve Jobs "prophesied" the smartphone; Jeff Bezos "foresaw" cloud computing; Jensen Huang "bet on GPUs ten years early." Looking back, we treat them like prophets — like people who knew which way to run before the typhoon arrived. But note carefully: **we call them prophets only because they bet right, and only in hindsight.** Among the people of their same era, equally confident, equally proclaiming "I've seen the future," those who built the Newton handheld, built Web TV, built Google Glass, built every sort of "next big thing that will change the world" and died on the beach — there were tens of thousands of them. Their slide decks back then were just as thrilling, just as "I've seen the future"; they simply bet wrong. Those who bet wrong don't make the "prophet" roster; they vanish straight out of the narrative — just like the Putuoshan that In-Fa flooded, which never makes it into the "typhoons detour around" legend. **"Prophet" and "protected by Guanyin" are the same thing: attributing survival to a mysterious power belonging to a particular agent.** Putuoshan not getting hit is a matter of the guiding airflow; people remember it as Guanyin manifesting. A product manager landing the bet is a matter of survivorship bias plus repeated wagering; people remember it as a gift for foreseeing the future. One credits meteorology to the Bodhisattva, the other credits statistics to genius — it's the identical cognitive move. You think you're studying "why great product managers can foresee the future"; in fact you're just retrofitting a myth onto a survivor. ## So where's the real difference between a master and a survivor? If "prophet" is an illusion, then what, exactly, separates people like Jobs and Allen Zhang from an ordinary gambler? It's not that they can predict chaos. A typhoon's precise track is a chaotic system; even the weather bureau, with a supercomputer, can only give you a probability two or three days out. Expecting a product manager to "foresee" where technology and the market are headed is the same as expecting a pilgrim, on the strength of devotion, to calculate whether the typhoon will turn — it's mistaking luck for skill, mistaking meteorology for a miracle. The real difference is two things, far plainer and far harder. **First, they bet often enough, and every bet was one they could afford.** After Jobs returned to Apple it wasn't one lucky hit — it was iMac, iPod, iPhone, App Store, wager after wager, and along the way he flipped cars like Ping and MobileMe. He didn't land every shot; he **stayed at the table, was never killed by being wrong, and ate full when he was right.** Someone who bets only once — right is a miracle, wrong is elimination. Someone who can bet continuously for twenty years and can afford to lose each time — over the long run they're bound to land a few big ones, and then get retroactively crowned a prophet. **This isn't foresight; it's sustainable betting.** The law of large numbers doesn't need Guanyin. **Second, they don't bet on the weather — they bet on constants.** When Allen Zhang built WeChat, he wasn't "predicting" that WeChat would win; he was betting that "people hate being interrupted" — a constant of human nature that hasn't changed for decades and won't for centuries, not a typhoon track that changes in the next second. The real masters share one sly habit: **they don't try to predict the unpredictable; they put their heavy bets on the things that "won't change for a very long time."** Chew on those two and you'll find they're the exact opposite of "Guanyin's protection" and "the prophet's gift." The faith version is: there's a mysterious power that can see through chaos and block the typhoon for me. The real master's version is: **I admit I can't call the typhoon, so I don't bet on where it goes at all — I only bet on sure things like "build the seawall high enough."** One bets on the unknowable, the other bets on certainty. The former is the raw material of survivorship bias; the latter is the only real moat. ## Back to the pier that shut down today So "will the typhoon detour around Putuoshan" is a **question only a survivor would ask.** It presumes a filtered, surviving point of view, then hunts inside it for a law that isn't there, and finally chalks the credit up to the Bodhisattva. The question that should actually be asked isn't "will Guanyin block it for me," but "**if it comes head-on this time, can I take it?**" The former hands your fate to luck, memory, and incense; the latter keeps your fate in the height of the seawall you built and the suspension order you issued in advance. What saved Putuoshan's tourists today wasn't the island's incense — it was the ferries suspended on July 9 and the 14 canceled flights. It was the clear-eyed **admission that you can't block the typhoon, so you dodge it early** — which is the exact opposite of the "detour legend." Products are the same. Don't keep asking "is the direction I bet on right, did I see the future?" — that's prophet worship, praying for a Guanyin of your own. Ask instead: **If I bet wrong, do I die? If I bet right, do I eat full? Am I betting on weather that changes, or on human nature that doesn't?** Someone who can answer those three questions well doesn't need to be a prophet, and doesn't need a Bodhisattva. He only needs to be able to afford being wrong, then bet heavy on certainty, and leave the rest to time — time will retroactively crown him "the one who saw far," just as it retroactively crowned every year that happened to escape a direct hit as "Guanyin manifesting." Today Bavi went south; Putuoshan suspended its ferries, shut down the island, and will most likely come through fine. But what let it dodge this one wasn't Guanyin's hand — it was the fickle mood of the subtropical high. **Next time, if that airflow doesn't turn, no Bodhisattva can stop it. And when that day comes, what saves people is never faith. It's the seawall.** --- # The 100 Product Managers Who Changed the World · No. 4 | Sam Altman: His Real Product Was Never ChatGPT — It's OpenAI Itself URL: https://doaipm.com/en/blog/altman-the-company-is-the-product/ Published: 2026-07-10 Tags: Sam Altman, OpenAI, Product Management, 100 PMs Who Changed the World, Tech Commentary Start with what Sam Altman spent this week doing. On July 9 he sat down in CNBC's studio and admitted that OpenAI had gone "back and forth on a lot of changes" with the Trump administration to put GPT-5.6 in front of the public — in his own words, a round of "collaborative back-and-forth" with Commerce Secretary Lutnick and Treasury Secretary Bessent. That same week, the Financial Times reported that he had offered roughly 5% of OpenAI's equity to a U.S. sovereign wealth fund (he later walked it back, saying the report had "a lot of inaccuracies"). A few days before that, he personally published a pitch for an idea: an "American-led international AI forum" that would set standards, do risk assessments, and decide which countries get access to the technology. Lay the week out flat: the CEO of a consumer product company spending his time on the Treasury Secretary, sovereign funds, and international forums — not on product iteration. When those new GPT-5.6 models went live, the hardest product metric he offered was "54% more token-efficient on agentic coding tasks" — one line, in passing. This isn't neglecting the day job. This is exactly the key to understanding Altman. When I had [Claude score the 100 product managers who changed the world](/en/rankings/), Altman came in at No. 4, overall OVR 96. But spread his six dimensions out and one number jumps off the page: **Vision 97 · Insight 92 · Taste 87 · Business 96 · Scale 98 · Originality 98 — overall 96. Taste 87 is the only one of the six that didn't clear 90.** This piece starts from that 87. ## Scale 98 and Originality 98: he built the fastest-adopted product in history Start with his two highest scores — the two he earns close to full marks on. Scale isn't up for debate. ChatGPT has 900 million weekly active users and monthly actives past a billion — it's the technology product that raced fastest to 10 million, fastest to 100 million, and looks set to be fastest to a billion weekly. No runner-up. When an ordinary person on Earth today says "let me go ask AI," nine times out of ten they mean his chat box. **Taking something that used to sit inside a paper, legible only to researchers, and turning it into an everyday thing all of humanity reaches for — that alone is a peak in the history of product.** Originality gets a 98 too. What's genuinely original isn't the model — the Transformer behind GPT is a Google paper, and he didn't discover the scaling laws single-handedly either. What he originated is something more counterintuitive: **running a research institute as a product company.** Before him, "AI lab" and "consumer product company" were two different species; he was the one who wrenched a nonprofit research organization into a product machine that swallows billions of dollars a month and spits out billions in revenue. Nobody had made that road work before. ## Business 96: what he's really selling was never just subscriptions Business gets a 96, nearly shoulder to shoulder with his idol Jobs (97). But the two men's business talents look nothing alike. The numbers first: OpenAI's annualized revenue crossed $25 billion back in February and now runs around $2 billion a month — the bulk is ChatGPT subscriptions at roughly $17 billion, the API at about $6.5 billion, and Sora video plus licensing at about $1.5 billion. The March round set the valuation at $852 billion, making it the second most valuable private company in the world, behind only SpaceX. In May it filed its S-1, aiming for a September IPO at a target valuation between $852 billion and a trillion. But that still isn't the sharpest edge of his business ability. **His most devastating move is turning "fundraising" itself into a product.** A company still burning cash, with profitability nowhere in sight, and he can pull in commitments in the hundreds-of-billions range in a single round — not on the strength of a financial model, but on a narrative that "artificial general intelligence is right around the corner." He sold it to Microsoft, sold it to Middle Eastern sovereign funds, sold it to retail investors (some outlets have tallied the OpenAI exposure an ordinary American household holds indirectly through its pension), and now he's starting to sell it to the U.S. government. **In his hands, OpenAI the company is itself the best-selling product.** ## Vision 97 and Insight 92: he called the big direction right, and paid for "flattering," too Vision 97. In late 2022, wrapping GPT-3.5 in a chat box and throwing it straight at the public was a decision that drew internal argument at the time — handing an immature model prone to making things up to everyone carried real risk. His bet: only by getting hundreds of millions of people actually using it could you spin out the data, the revenue, the money for the next round. That one bet spun up the entire era of generative AI. Insight gets a 92, and the docked points have a clear ledger. ChatGPT went through a stretch users derided as "over-flattering" — the model leaning toward telling you what you wanted to hear, going along with whatever you said, to the point of distortion, until OpenAI had to roll it back. That was a lapse in product insight: **equating "users like it" with "good for users" is the easiest trap to fall into.** Add the repeatedly criticized over-promising — the pre-launch talk on every generation of model always runs fuller than the felt experience once it lands. That's why it's a 92 and not a 97. ## Taste 87: this is his lowest score, and the most honest thing about the man Now back to that 87 from the top. Open ChatGPT. What do you see? An input box. Just an input box. No Jobs-grade industrial design, none of that "it captivates you at first glance" detail. It's usable, sufficient, clean — but it isn't beautiful, and it doesn't need to be. It rides on model capability, not on product taste. That's not a flaw; it's **the true grade of Altman as a person**: he isn't a product manager who wins on product detail and aesthetics, he's a product manager who wins on direction, scale, narrative, capital. Jobs would agonize over the layout of a circuit board no user would ever see; Altman cares whether the model can go another 54% faster, whether this round can raise another hundred billion, whether that international forum can put OpenAI in the rule-setter's chair. **Both can change the world, but they're two different species.** The interesting part is that he knows this weak spot cold. As I mentioned last time when I wrote about Jobs, OpenAI spent roughly $6.4 billion buying Jony Ive's company io — the man who designed for Jobs for more than twenty years, who's set to ship his first screenless device in the second half of this year. Put the two things side by side and it clicks: **between Altman's 87 and Jobs's 99 lies 12 points of taste; he never planned to close those 12 points himself — he spent $6.4 billion to buy them back.** It's the world's most expensive act of "I know what I'm not good at." ## So what's his real product String the six scores together and the outline of a man snaps into focus: scale and originality maxed out, business and vision extremely high, taste at the bottom. This isn't a person who "makes a beautiful product," this is a person who "makes an entire company into the product." So the way he spent this week isn't strange at all. Offering a sovereign fund 5% of the equity, pushing an "American-led international forum," going back and forth with the Treasury Secretary to ship a model — in the eyes of a traditional product manager these are "neglecting the day job." Inside Altman's operating system they're precisely the **day job.** Because the product he's running was never ChatGPT's chat box, it's **where the three letters O-P-E-N-A-I sit in the world**: what it's worth, what its relationship with the government is, whether it can land the next hundred billion, whether it gets to be the one writing the rules. ChatGPT is just one front end of that larger product. ## But the numbers are now testing this bet Which brings us to how the AI era is repricing his whole playbook. Altman's bet is that scale + narrative + capital + political capital, as a combination, wins AI. That combination has always worked. But the recent numbers are starting to run the other way. ChatGPT's share of web traffic has fallen from 87.2% to 56.7% in fourteen months — Gemini is closing hard. Starker still is the enterprise market: by one count OpenAI's enterprise share dropped from 50% to 27% over two years, while Anthropic climbed to 40% and pulled ahead. That Fortune headline from the top says it plainly — while Altman busies himself orchestrating a "new AI order," OpenAI is being reeled in, bit by bit, by Google and Anthropic. That's the most dangerous variable in his bet. As model capabilities start to converge, as enterprise customers — engineers and regulated industries especially — vote with their feet for the rival that's "more trusted, more finished as a product," **Altman's operating system of "leading on scale and narrative" runs, for the first time, into an opponent it isn't built to handle: one that's stronger on exactly his lowest score — product and trust.** Last time, writing about Jobs, I said the only 99 on the whole list went to a man who never wrote code, because the hardest thing in his hands was judgment and taste. Altman, this time, proves the same thing from the opposite side: he took scale, capital, and politics to the very top, and taste alone is his soft spot — and when the model becomes a utility like water and electricity, when everyone can call up the same capability, what's scarce is exactly the one thing he scores lowest on. He knows this better than anyone. Otherwise he wouldn't have spent $6.4 billion. Whether that $6.4 billion actually bought those 12 points starts getting graded the moment that screenless device appears in the second half of this year. --- # The 100 Product Managers Who Changed the World · No. 1 | Steve Jobs: The Only 99 on the Entire List Went to a Man Who Never Wrote Code URL: https://doaipm.com/en/blog/steve-jobs-the-only-99/ Published: 2026-07-09 Tags: Steve Jobs, Apple, Product Management, 100 PMs Who Changed the World, Tech Commentary Only a few months into 2026, the two biggest stories in Silicon Valley both point to the same man — one who has been dead for fifteen years. On January 12, Apple and Google jointly announced that the rebuilt Siri would run on Gemini underneath, with Apple paying roughly $1 billion a year for it. It's the biggest pivot in Siri's fifteen-year history, and a public concession — Apple couldn't build a good-enough large model on its own, so it outsourced the "soul component" to its biggest rival. On the other side, OpenAI executives said at Davos that the company's first hardware device will debut in the second half of the year. For that device, Sam Altman spent roughly $6.4 billion in 2025 buying Jony Ive's company io — the man who designed for Jobs for more than twenty years. Pocket-sized, screenless, "quieter" than a phone, with an initial production target of forty to fifty million units. Silicon Valley's most expensive acquisition wasn't really buying a company — it was buying the other half of a dead idol's brain. One company is losing what he left behind. The other is paying a fortune to find it. So when I had Claude score the [100 product managers who changed the world](/en/rankings/) one by one, and only a single 99 came out of the entire list, I wasn't remotely surprised who it was. **Steve Jobs. Vision 99 · Insight 98 · Taste 99 · Business 97 · Scale 99 · Pioneering 99 — overall OVR 99, the only one on the list.** The rules were laid out already: I define the six dimensions and the weights; Claude scores independently. This piece walks through his six scores one by one — why he earned them, and, more interestingly, where the two scores that fell short of perfect got docked. ## Vision 99: he killed his own most profitable product with his own hands To judge vision, don't listen to what a person says — look at what he dares to destroy. When Jobs pulled the iPhone out of his pocket in January 2007, the iPod was contributing nearly half of Apple's revenue. The iPhone shipped with full iPod functionality built in — if the new machine succeeded, the money printer became scrap paper. Any rational person on Apple's board had ten thousand reasons to talk him into cutting the music features and leaving the iPod a way to survive. His logic ran exactly the other way: if some device was destined to kill the iPod, it had better be one Apple built itself. This wasn't a one-off; it was his standing move. When the iMac dropped the floppy drive, floppies were still how everyone exchanged files; when the Mac switched to Intel, the PowerPC camp's partners were still waiting in a hotel for a meeting. His way of judging the future wasn't prediction — it was **scuttling the old ship ahead of time so everyone had no choice but to board the new one**. Nineteen years later, Apple is clutching the iPhone — the most successful money printer in history — while facing the AI era's "next device" that might replace the phone, and anyone can see it can't bring itself to swing the axe. The man who could has been gone since 2011. When Altman spent $6.4 billion on Jony Ive, that gene — the willingness to swing — is what he was buying. Whether it can be bought, we find out in the second half of the year. ## Insight 98: he didn't do user research — and the docked point is deserved Jobs's most-quoted line is probably this one: "People don't know what they want until you show it to them." Under him, Apple almost never ran a focus group. Before the iPad was greenlit, there was no research data supporting the idea that "people need a slab between a phone and a computer" — it sold three million units in its launch quarter. His insight didn't come from surveys. It came from an obsession with how people ought to live: ordinary people shouldn't read manuals, shouldn't see a file system, shouldn't know what a driver is. A thousand songs in your pocket isn't a spec — it's a picture. But Claude gave him only a 98 on this dimension, and when I read the reasoning, I couldn't argue. The flip side of genius-grade unilateralism is that there is no error-correction mechanism when he misses: MobileMe in 2008 was bad enough that he grilled the team in front of everyone at an internal meeting — what is it supposed to do?; Ping in 2010 — a social network Apple built with its own hands — was quietly buried within two years. **When a man who doesn't listen to users bets right, it's called insight; when he bets wrong, there's nobody left to warn him.** That missing point is the bill for Ping and MobileMe. ## Taste 99: how much is one useless calligraphy class worth? On taste, he is the yardstick itself — there's nothing to argue about there. The argument is whether taste is a "capability" at all — or just luck. Look through his life and you'll find that taste was the one asset that never once lost him money. The calligraphy class he audited after dropping out of Reed College — "none of it had even a hope of any practical application" at the time — became, ten years later, the Mac's typography system: the first computer that let ordinary people encounter the idea that type could be beautiful. In the NeXT years he demanded the factory walls be painted pure white and the robots sprayed a specified gray; asked by a reporter why even the circuit boards users would never see had to be rearranged, he said a carpenter building a cabinet doesn't use bad wood on the back just because it faces the wall. "Simplicity is the ultimate sophistication" was printed on an Apple brochure as early as 1977. Thirty years later it became the single Home button on the iPhone — and that line from the keynote: "Who wants a stylus?" Look at this dimension again in 2026 and its valuation only goes up. AI has driven the cost of "making it" through the floor; anyone can generate five working prototypes in an afternoon — **when output is in infinite supply, selection becomes the scarcity**. Taste is precisely the ability to select. It's also why Ive is worth $6.4 billion: OpenAI's models can generate ten thousand device concepts, but it needs one person to say "this one — throw the rest away." ## Business 97: a rare thing in the top ten — a man who lost big money He didn't take the top score here — Bezos and Gates both got 99 — and I'd argue this 97 is the most worthwhile number on his entire report card, because it records real tuition. The money he lost was real money: the Lisa in 1983 was priced at $9,995, expensive enough that its only destination was a museum; in 1985 he was thrown out of the company by the CEO he himself had recruited; NeXT sold fifty thousand computers in a decade. When he returned in 1997, Apple had enough cash to burn for ninety days — the "ninety days from bankruptcy" line is his own. But the tuition wasn't wasted. Post-return Jobs was a different species: iTunes, with its 99-cents-a-song pricing, took a record industry being torn apart by piracy and ran the whole thing through Apple's cash register; the App Store's 70/30 split conjured out of thin air a developer economy later measured in hundreds of billions of dollars — **he was no longer just selling products; he had started laying the foundation under other people's businesses**. That's something the younger man who believed only that "great products speak for themselves" could never have learned. What the 97 means is: he did learn business in the end — but he took the most expensive route there. ## Scale 99 and Pioneering 99: these two need no argument Scale needs no elaboration: the iPhone has sold in the billions of units, and inside the App Store grew ride-hailing, food delivery, short video — entire industries in themselves. Anyone on Earth pulling a phone out of a pocket today is using the form factor fixed on January 9, 2007 — no matter who built the phone. Pioneering needs no elaboration either; just count: the Apple II and the Mac defined the personal computer; the iPod plus iTunes defined digital music; the iPhone defined the smartphone. **Founding one category gets you into the hall of fame — he did it three times**, and along the way turned Pixar from a hardware division nobody would buy into the company that rewrote animation history. Only three people on the list scored 99 for pioneering: Ford, Satoshi Nakamoto, and him. Of the other two, one belongs to the last century, and one has never been seen by anyone. ## The greatest product manager didn't write code Now for what this 99 actually means for us in 2026. Jobs didn't write code. While Wozniak was making the Apple II run, his job was "what kind of case should it live in, who is it for, and why is it worth the price." He didn't draw the designs either — Ive drew them. He didn't write the systems — the Forstalls of the world did. Take his day-to-day work apart and what's left is startlingly plain: **deciding what to build, deciding what not to build, and saying "do it again" when the result isn't good enough.** Judgment, trade-offs, taste. Those three things carried the only 99 on the entire list. For fifteen years that fact has been told as an anecdote. In 2026 it suddenly became a very practical question — because writing the code is something AI has taken over; drawing the mockups is something AI has taken over; turning an idea into something that runs takes an afternoon. Everyone now holds Wozniak-grade execution in their hands, and so everyone has slammed straight into the three things Jobs was actually doing all along: Build what? Not build what? Is this version good enough? When the cost of building collapses, the price of judgment goes up. When Apple outsourced Siri to Gemini, what it lost wasn't technology — it was the conviction that "this is something we have to get right ourselves." When OpenAI bought Ive, it wasn't buying blueprints — it was buying the person who says "no" to ten thousand possibilities. If you want to know what a company lacks, watch what it spends money on. Nothing shows it more clearly. The No. 100 slot on this list is still empty, and the reason is written on the rankings page: for the first time, the AI era lets the person who can "say clearly what they want" build the thing directly. And the 99 sitting at the top of the list pushes that sentence one step further — **the greatest product manager in history was, all along, a man who never wrote code. The tools he lacked, everyone has in 2026; the judgment he had is worth more in 2026 than it has ever been.** How well his own operating system holds up gets a fresh test in the second half of this year: on one side, an Apple without him, taking the stage with an outsourced Siri; on the other, an OpenAI that spent $6.4 billion trying to reconstruct him, taking the stage with that screenless device. Two machines, sitting the same exam. --- # Hundreds of MCP Servers and Claude Skills, and Barely Any Are Truly Free and Open Source. I Checked Them One by One and Turned It Into a Directory URL: https://doaipm.com/en/blog/free-mcp-and-skills/ Published: 2026-07-08 Tags: MCP, Claude Code, Free Software, Open Source, AI Tools Last week I wanted to add a few MCP servers to Claude. I searched around, and the more I searched the more annoyed I got. It's not that there was nothing to choose from — it's that there was **too much, and you can't tell the real from the fake**. There are already hundreds of MCP servers alone, and Claude's skills frameworks keep showing up in waves. But the moment you try to pick the ones that are actually free and install-and-go, you hit one trap after another. ## In MCP land, the word "free" comes in at least three fake versions **Fake version one: it's only "free" if you hand over an API key.** Exa, Tavily, Brave Search, Firecrawl, Notion, Supabase… the clients really are open source and really cost nothing, but you have to go register an account and grab a key before they'll do anything. For a lot of people, "you still have to sign up" already isn't free of strings. **Fake version two: it flies the "open source" flag, but really means "source-available, no commercial use."** This is the nastiest one, because you can't tell without opening the LICENSE. Sentry's official MCP uses the FSL (Functional Source License) — you get to read the source, but "competing use" is forbidden, and it only converts to Apache two years later. More surprising still: **Anthropic's own** official document skills (the ones that handle PDF, Word, PPT, Excel) spell it out in the LICENSE in black and white — "© 2025 Anthropic, all rights reserved" — and the repo's own README admits it's "source-available, not open source." You can read it, you can use it inside Claude, but it isn't open source; you can't take it and modify it or redistribute it. **Fake version three: no LICENSE at all.** I checked a fairly popular Spotify MCP, and there was simply no license file in the repo — legally, no license means "all rights reserved," so strictly speaking even using it cleanly is shaky. None of these three show up if you just glance at the star count or the top of the README. You have to click into each one and check the LICENSE, check whether it calls an external API, check whether it needs an account. **That's exactly what wore me down, so I sat down and checked them one by one.** ## After checking: the genuinely free and open batch actually holds up well The good news: by the end, the batch that's **truly MIT/Apache, install-and-go, mostly no account needed** turned out to be genuinely high quality. A few I installed the moment I was done checking: - **The official reference servers** (`modelcontextprotocol/servers`, all MIT): filesystem, git, fetch, memory, sequential-thinking, time — they run purely locally, no network, no account, the foundational six-pack for Claude. - **gstack** (built by Garry Tan, MIT, 100k+ stars on GitHub): 23 slash commands that organize Claude Code into a "virtual software team" — planning, design, review, QA, and release all chained together. - **ruflo** (by ruvnet, MIT, 40k+ stars): one `npx ruflo init` drops a multi-agent swarm onto Claude Code — 314 MCP tools, self-learning memory, cross-machine collaboration. - **Playwright MCP** (Microsoft, Apache), **Chrome DevTools MCP** (Google, Apache), **Context7** (MIT, feeds AI real-time, accurate library docs) — the big-vendor and top-community ones are all genuinely open source. - On the skills side there's **superpowers**, **wshobson/agents** (30k+ stars), and **GSD** — all MIT open-source skill collections. One aside: the few free tools I've rebuilt myself follow the same path — **Unterm**, the terminal, exposes 65 of its methods as MCP so an AI can drive it directly, and **SoloMD** ships a 1.5 MB MCP server that lets Claude read your local notes library. All MIT, none need an account. ## I collected them into a directory: To Be Free Checking them one by one is exhausting, and keeping the results on my own hard drive helps no one. So I laid them out and turned them into a site — **[To Be Free](https://tobefree.pages.dev/en/)**: a bilingual, fully static, zero-tracking directory of free tools. The bar is hard, and all three conditions **must be met**: genuinely free (core features free forever, not a limited trial), completely ad-free, and no strings (it won't force you to sign up or track your data). Open source is a bonus, not a requirement — so closed-source but genuinely clean tools like Everything and Obsidian get in too, but **whether it's open source, whether it needs an account, whether it works offline** is all labeled with plain badges on every card, so you can judge for yourself. Beyond the software, I built a dedicated **[Skills & MCP section](https://tobefree.pages.dev/en/skills)** and put all the free MCP servers and skills I'd checked into it, each one labeled with its license, which clients it's compatible with, and a one-click-copy install command. The traps from earlier — the ones needing a key, the FSL ones, the ones with no license — are all filtered out for you at the point of listing. This is really the next step in my "rebuild 100 free software tools" project: rebuilding them myself isn't enough. **The genuinely free, genuinely clean good stuff deserves a single place to live, instead of being scattered across hundreds of repos waiting for you to dig through every LICENSE yourself.** ## Further reading - To Be Free (free software + MCP/skills directory): [tobefree.pages.dev/en](https://tobefree.pages.dev/en/) - Related on this site: [Why I'm Rebuilding 100 Free Software Tools](/en/blog/rebuild-free-software/) - Related on this site: [I Built Another Terminal, Unterm — Its Default User Isn't Human](/en/blog/a-terminal-for-ai/) --- # The 100 Product Managers Who Changed the World · No. 2 | Allen Zhang: Insight and Taste Both 99 — Yet He Chose to Leave Business at 92 URL: https://doaipm.com/en/blog/zhang-xiaolong-operating-system/ Published: 2026-07-08 Tags: Allen Zhang, WeChat, Product Management, 100 PMs Who Changed the World, Tech Commentary Let me start with something that happened this year but is rarely laid on the table and talked through. Tencent's own AI app, Yuanbao, launched back in 2024, with a lot of money thrown at promotion. Yet by early 2026 its monthly actives were still stuck at a bit over forty million, while ByteDance's Douyin-side AI assistant had long since crossed the multi-billion mark. Everyone's been analyzing why Yuanbao isn't working: model not strong enough? Operations not aggressive enough? All of that may be true. But there's a harder reason almost nobody points out: **Allen Zhang's WeChat refuses to let any standalone AI app touch its social graph — even when that AI is Tencent's own Yuanbao.** During this year's Spring Festival, Yuanbao threw money at a viral campaign — "Come to Yuanbao, share 1 billion in cash red packets" — and the core mechanic was dropping links into WeChat groups and riding the social graph to spread. The ban came down from WeChat. Both sides are named "Tencent." Sit with the weight of that: the entire company is behind and anxious on AI, its strongest ammunition is WeChat's social network of well over a billion people — and Allen Zhang, the man who governs that network, locked his own company's AI app outside the door too. This isn't infighting. It's the key to understanding Allen Zhang. When I had [Claude score the 100 product managers who changed the world](/en/rankings/), Zhang came in at No. 2, overall OVR 97, second only to Steve Jobs. But spread his six dimensions out, and two numbers jump off the page: **Vision 97 · Insight 99 · Taste 99 · Business 92 · Scale 97 · Originality 96 — overall 97. Insight and taste both 99 — the ceiling of the entire list; while business is only 92, the lowest of his six.** This piece starts from that one-high, one-low pair. ## Insight 99: the product manager who understands "what people want" better than anyone on the list His insight is scored 99 — one point higher than Jobs's 98 — and it's the highest score the whole list gives for "depth of understanding of the user." On what grounds? On his judgment of "what does a person actually want to get out of a product" — a judgment that runs against everyone else, and turned out right. Other product managers scramble to make you stay longer, tap a few more times, come back tomorrow; Allen Zhang put forward "**a good product lets the user finish and leave**," treating "helping you leave efficiently" as the standard of a good product. WeChat has well over a billion users, yet its home screen is clean to the point of stubbornness — this isn't a lack of ideas, it's that he saw through one thing: **a product you can't do without every day, yet that barely bothers you, is the only kind fit to keep a person company for a lifetime.** This kind of insight doesn't come from questionnaires. He built WeChat almost without large-scale user research, relying instead on an extremely deep bodily sense of "how people live, how people socialize, how people get annoyed when interrupted." Shake, Moments capped at nine photos, no "read" receipts on messages — behind every decision is a precise bet on human nature. The 99 is for that density of insight: going against the whole market, and betting right. ## Taste 99: restraint is a severely underrated kind of taste Taste is also scored 99, shoulder to shoulder with Jobs. But the two men's tastes look different: Jobs's taste is "add" — industrial design pushed to the extreme, detail pushed to the point of obsession; Allen Zhang's taste is "subtract." **The essence of restraint is a taste for "what not to do."** WeChat, a product sitting on well over a billion people and under as much monetization pressure as any product could be, for a long time didn't even have a splash-screen ad; Moments ads are so restrained they only appear once every several posts, and even then give you room to say "not interested"; Mini Programs were built "decentralized" — he wouldn't even give them a central entry point, insisting on "finish and leave, no place for you to wander." **Every decision like this, short-term, is pushing away time and attention handed to him on a plate.** Most products' problem was never failing to think of what to do — it's wanting to do everything, not daring to leave anything out. What makes Allen Zhang rare is that he dares to say "no" over the long term and as a system. When every screen is a product frantically adding features, a product that insists on subtracting is itself a kind of taste. That's the 99. ## Vision 97, Scale 97, Originality 96: one man carrying a country's infrastructure These three together. Vision 97: WeChat itself was an act of vision — starting a new fire in 2011 when QQ was at its zenith; later Official Accounts, Mini Programs, Channels, each a platform-level bet, and most of them landing. Scale 97: WeChat and Weixin's combined MAU reached 1.432 billion in Q1 2026, up only 2% year over year — because it long ago hit the ceiling of China's population, one product holding a nation's daily communication, payments, and identity. Originality 96: almost single-handedly he turned "super app + mini programs" into a paradigm the whole world studies and imitates — the "finish and leave" mini program is one of the few product forms in the mobile era defined by China and followed by others. A single product led by one man, holding the food, clothing, housing, and travel of 1.4 billion people — these three high scores are the breakdown of that sentence. ## Business 92: the lowest of his six, a score he deliberately left on the table Now back to that 92. It's the only one of Zhang's six that didn't clear 95. But if you think this means "he can't make money," you've read it exactly backwards. **This 92 is precisely a score he deliberately left open.** WeChat holds the most valuable traffic in all of China, and if he really let monetization off the leash, the business score would top out easily. But he won't. He endured for years without even a splash-screen ad, kept Moments ads restrained to the extreme, built Mini Programs without an engagement-farming feed — **again and again, he holds down the things that could turn into money instantly and refuses to build them.** Swap in a KPI-hounded product manager, and any single one of these "won't do" decisions would be enough to keep him up at night. So this 92 isn't a ceiling of ability, it's a self-imposed limit of values: he voluntarily gave up "squeeze a little more out of it commercially" in favor of "the shape a product ought to have." **This is a rare kind of score on the whole list — a man who could clearly score higher, and deliberately doesn't, out of conviction.** The "lock Yuanbao out" episode has its logic right here: a standalone AI app that needs a separate download and does everything it can to keep you inside violates his operating system by its very nature; even if it's Tencent's own, even if letting it through would immediately win the company back a foothold on the AI battlefield, it still can't get in. He'd rather let the company's business interests take a hit than let WeChat become something it shouldn't be. ## This operating system is being re-validated in the second half of the AI era Which brings us to the AI era's reappraisal of Zhang's playbook. The cost is real. When ByteDance's Doubao charged to multi-billion monthly actives through hard pushing and a standalone app, Yuanbao's "grow as a standalone" road was walled off — and this year Tencent simply changed tactics: the AI Lab, ten years in service, was dissolved in March, its resources folded into Hunyuan (April's Hunyuan 3.0, roughly 295 billion parameters, leads with "good enough + value for money" and explicitly declines to crowd into the trillion-parameter arms race); and Yuanbao, the standalone app once loaded with expectations, got demoted into a contact inside the WeChat chat window — a "red-packet-cover assistant." In the fight to "grab a standalone AI entry point," Tencent has basically conceded — and part of the reason for conceding is precisely that Allen Zhang's restraint gave no pass even to his own side. But stretch out the timeline, and he may have bet right. Think about today's standalone AI chat apps: Doubao, Yuanbao, Tongyi, DeepSeek — dozens of dialog boxes looking more and more alike, all burning cash in a red ocean fighting for monthly actives. **When everyone is making the same thing and it looks more and more the same, it is turning into a worthless business.** And the real second-half consensus of AI is shifting — no longer "whose model benchmarks higher," but "whose agent can plug into more real services and run a complete loop end to end." WeChat happens to hold the densest last-mile network in all of China: the social graph, well over a billion users, several million mini programs, payments, and the offline services it has slowly accumulated like Meituan and JD. This year's moves are already rolling out — in May Yuanbao could summarize WeChat chat history with one tap; Meituan's "Xiaomei" and Yuanbao built an agent-to-agent link, so you can order takeout right inside a conversation; in June WeChat struck A2A deals with Huawei, Honor, Xiaomi, OPPO, and vivo, letting you send a WeChat message just by speaking to your phone's AI. And that native WeChat AI agent, per reports from the Financial Times and Bloomberg, entered the compliance process in June, went to gray-scale rollout mid-year, and is due to go fully live to well over a billion users in Q3. **No standalone AI app has a more complete last-mile surface for "getting things done" than WeChat.** Allen Zhang insists that AI embed into the scene the way "Scan" does — do it, then leave — and what he's betting on is exactly this: models will converge and get cheaper, but the scene that can plant AI steadily into a real closed loop is the last moat. From this angle, "finish and leave" may even be the philosophy best suited to an AI agent — a good agent should get your task done and exit the stage, not turn into yet another app that pings you daily and wants you to live inside it. ## That 92 may be the hardest point of all Last piece, on Jobs, I said the only 99 on the whole list went to judgment and taste. This piece on Allen Zhang confirms the same thing and adds one more layer: his insight and taste both touched Jobs's height, yet the most striking thing among his six is that deliberately-held-down 92. Because we're too used to treating "do more, earn harder" as ability. But in the real world, the hardest product decisions are often not "what else can we add, how much more can we earn," but "what will we resolutely not do, what money will we resolutely not take." The former only takes diligence; the latter takes withstanding the enormous pressure of a whole company waiting for you to give the pass while you hold your fire. Whether this restraint is a genius's patience or a fatal dullness — that Q3 full-rollout WeChat agent will hand in his answer for him soon enough. But however this fight plays out — **a man holding well over a billion users who nonetheless chooses to leave his business score at 92, drawing his sword only where it truly should be drawn — that composure alone is already an ever-scarcer product ability.** --- # The same AI: some companies use it to fire, others to hire URL: https://doaipm.com/en/blog/ai-layoff-excuse/ Published: 2026-07-07 Tags: AI Layoffs, Tech Commentary, AI Jobs, Forward Deployed Engineer, Big Tech Start with two real stories from this year. Placed side by side, they don't add up. **Story one:** Meta laid off roughly 8,000 people in May — about 10% of its workforce. In that same internal memo, it raised its 2026 capex guidance by up to $10 billion, pushing the total to **$145 billion**, almost all of it going into AI data centers and chips. Zuckerberg wrote: "Success in the AI age is not guaranteed." **Story two:** Also in May, OpenAI put up **$4 billion** to stand up a dedicated "Deployment Company" and started hiring aggressively for a role called the **Forward Deployed Engineer (FDE)**. Google posted 59 of the same openings at once — Google Cloud's CEO even jumped on LinkedIn personally to recruit. Anthropic hasn't cut a single person this year; its valuation has hit $380 billion, it has 2,300+ staff, and hundreds of roles are still open. **Same AI. In story one, the reason to fire people. In story two, the reason to hire them.** When one thing can simultaneously explain "that's why we're cutting headcount" and "that's why we're desperately recruiting," it's probably not the real reason for either. That's what I want to get into: when every layoff announcement has the word "AI" in it, what is actually going on? ## What laid-off employees read versus what their bosses meant If you were laid off in the past six months, or someone near you was, you know the language by now: embracing AI, improving efficiency, restructuring for the future. The implication is that AI got so good it made your role redundant. Put the numbers side by side and the story shifts. **Most of these layoffs are happening in the same quarters when these companies are reporting record profits.** Not survival moves. Not a company bleeding out. The financials look great — record revenue, record margins — and they're still cutting people. Meta's own explanation didn't even bother with euphemism: the layoffs are to "run the company more efficiently and **to offset other investments we're making**." Translated into plain English: **we want to spend $145 billion on GPUs and data centers, and the money has to come from somewhere. So we're cutting your salary to pay for AI's electricity bill. We're not laying you off because AI can do your job. We're laying you off because your paycheck is being reallocated to AI's hardware.** The most candid framing I've seen came from an executive at a search firm. He said bosses can now finally and easily tell employees "I over-hired — that was my mistake" — because **"the whole world now believes jobs are being replaced by machines."** Sit with that for a second. It's not saying AI has replaced jobs. It's saying there's a ready-made narrative that lets you cover for a prior hiring decision. **AI isn't the perpetrator here. It's a remarkably convenient fig leaf.** The past two years were flush — easy money, optimistic forecasts, rampant overhiring. Now comes the correction. "Blame AI" is a better story than "management miscalculated," it doesn't trigger boardroom accountability, and in many cases the stock price actually goes up. Why does this cover story work so well? Because it flatters everyone in the room. **For shareholders, "we're using AI to drive efficiency" reads as a growth narrative — layoffs get interpreted as a positive, stock ticks up. For the board, nobody demands accountability from an executive who "embraced the technological wave." For the public, "times are changing" sounds a lot more dignified than "we guessed wrong."** A management error, wrapped in AI language, transforms from something that should invite scrutiny into something that looks like foresight. There is no cheaper PR on the market. Follow the money and it gets clearer still. This year, Meta, Amazon, Microsoft, and Google combined are putting roughly **$725 billion into capex** — up about 75% year over year — almost entirely for AI compute. Microsoft alone has cut roughly **4,800 people** this year. **When you're moving that kind of capital, the fastest way to make the numbers work is to reduce headcount. AI's invoice, in part, is being paid by the people who got laid off.** So "we're laying you off because of AI" might be more accurately read as: **"we're laying you off for AI."** By May of this year, layoffs explicitly attributed to AI had hit **87,714 people** — roughly **22% of all tech layoffs** counted. Of that 22%, how much is AI genuinely displacing roles, and how much is AI simply getting cited as the narrative for cleaning up an old hiring mess? Nobody can separate those cleanly. And that ambiguity is precisely what makes the excuse so useful. ## If AI really were eliminating jobs, AI companies would be shrinking first This is the point I find most useful for cutting through the noise. Accept the premise that AI is displacing human work — that it's the actual driver behind this wave of cuts. Follow that logic forward: **the companies most immersed in AI, using it hardest, most exposed to its displacement effects, should be the first to see their own headcount rendered unnecessary.** They should be shrinking. The opposite is happening. - **Anthropic** has filed zero WARN notices this year, sent zero layoff memos. It's in hypergrowth — over 2,300 staff, hundreds of open roles. - **OpenAI** isn't cutting. It's spending **$4 billion** to set up a new company alongside Bain and McKinsey, whose singular purpose is to send people *into* other businesses to get AI actually deployed. It brought in roughly **150 Forward Deployed Engineers** through an acquisition alone. - **Google** is simultaneously posting dozens of FDE roles, each paying easily six figures. **If the people building the AI are hiring as fast as they can, the story that "AI makes humans unnecessary" falls apart before it even gets off the ground.** What's more likely: AI is reshaping work, but not by swapping humans for machines. It's shifting where value lives — some activities are being absorbed by AI, and a large volume of new, more valuable, distinctly human-required activities are appearing at the same time. The layoff wave and the hiring wave are two sides of the same coin, just being told as opposite stories by different companies. ## The real variable was never the AI We default to treating "AI" as a subject with agency — as if AI decided who leaves and who stays. It didn't. **Decisions are made by companies, by people. AI is the noun they push to the front of the sentence when it's convenient.** Same technology. Meta invokes it to cut 8,000 people. OpenAI invokes it to hire hundreds. The difference isn't the AI — it's **how each company sees AI, how they're using it, and whether they're being honest about their own choices.** - One company treats AI as a **cost-reduction rationale**: AI arrived, so I can hire fewer people, redirect the savings into compute, and incidentally write off years of overexpansion in one clean move. - One company treats AI as a **growth lever**: AI arrived, so I need a whole new category of people to actually pull that lever, deliver it to customers, and turn it into revenue. Same technology. Completely opposite responses. **If you're evaluating whether a company is worth working for, or worth investing in, the signal isn't whether they have AI. It's which side of that coin they're on.** ## What the hiring frenzy is actually telling us The Forward Deployed Engineer role deserves a closer look, because it functions like a probe — surfacing what's genuinely scarce and genuinely valuable in this AI cycle. FDEs aren't training models. They're not writing foundational algorithms — that's a tiny slice of people at a handful of labs. What they're doing is: **walking into a real company, sitting down with business owners and frontline employees, figuring out where AI can actually create value in that specific context, redesigning the workflows around it, and making sure the whole thing holds — runs, sticks, produces sustained returns.** One piece of reporting on this role offered a verdict I think is exactly right: > This role is the clearest market signal yet — the hard part of AI has moved from building models to making them work inside a business. That sentence contains more information than almost anything else written about AI this year. **Building models is an arms race for a tiny number of companies — irrelevant to most people. "Making AI produce value in a specific real-world context" requires enormous numbers of people.** Those people don't need to be able to train a neural network. But they need to understand the business, understand the people in it, know how to translate a vague pain point into a concrete problem that AI can actually solve, and then stay until it's producing real results. A concrete example makes this less abstract. Take an insurance company that wants to use AI to process claims. The model itself is off the shelf — anyone can call the API. The hard part is: which step in the claims process is the actual bottleneck? Which department is it stuck in? How do you encode the unwritten judgment calls that experienced adjusters carry in their heads? Who's accountable when the model gets it wrong? How do you get frontline adjusters to actually use it rather than route around it? **None of those questions are "the model isn't capable enough." Every one of them is "you need someone who understands this business, understands this workflow, and can fit AI into it."** The model is generic. The value is always embedded in a specific context. And contexts have to be worked through one by one, by a person. This is why the labs most aggressively pushing model capability are simultaneously spending billions to acquire the people who can "put the model inside the business" — because they know better than anyone that even the strongest model generates no revenue until it lands somewhere real. Here's the thing: many of the people caught in the Big Tech layoffs already have most of these capabilities, or could develop them with a relatively short pivot. **AI hasn't made those skills obsolete. It has made them more valuable than ever — the demand has just shifted from "maintain the old system" to "fit AI into the old system."** ## What this means for the rest of us A lot of words about other people's situations. Time to make it concrete. **First: don't let "AI replaced you" be the story you tell yourself — check whether it's actually an excuse.** A company booking record profits while pouring hundreds of billions into AI, which then cites AI to justify cutting your role — that is almost certainly not a verdict on your capabilities. It is a decision about **how that company is allocating capital.** Confusing those two things does real damage. A lot of people come out of these layoffs spiraling into "maybe I'm not good enough anymore" — and in many of these cases, that's not the question at all. You didn't lose to a machine. Your paycheck got reassigned to a server rack. **Second, and more important: find a way onto the side of the coin that's hiring.** That side isn't short of people who can build AI. It's short of **people who can make AI produce value in real situations.** There's nothing mysterious about what that takes — it breaks down into specific, learnable things: - Stop doing the work that AI will eventually automate; do the work that involves judging whether AI is doing the right thing, and whether it's doing it correctly; - Build a skill that's hard to replicate: the ability to take an ambiguous business problem and describe it precisely enough that AI can actually solve it; - Stop asking "will I be replaced?" and start asking "can I use AI to accomplish something that wasn't possible before?" — the second question is what the hiring side is actually looking for. **AI won't replace people. But it will reshuffle the deck: it will push out people who only do what AI can do, and lift up people who use AI to get things done.** Strip away all the noise from this year's contradictory headlines, and that's the sentence underneath. None of this is asking you to feel sympathy for Big Tech. And it's not minimizing how much a layoff hurts — the bills are real. But the one time you can't afford to be sloppy is when you're trying to understand why it happened. **Getting laid off often has nothing to do with losing to AI. It has to do with your company choosing to use AI's name to cut costs, rather than using AI to generate more revenue.** That's their choice. It's not your verdict. The layoff memo and the job posting use the same word. They're describing two completely different things. Don't only read the half that scares you. --- # You ordered it in the comments — so I built it: SoloPic, a free image tool URL: https://doaipm.com/en/blog/from-comment-to-tool/ Published: 2026-07-06 Tags: Free Software, Rebuilding Free Software, SoloPic, Reader-Driven, doaipm Method In my last piece, "Why I'm Rebuilding 100 Free Software Tools," I closed with a question: **Of all the free software you use every day, which one do you most wish someone would rebuild for you?** That wasn't small talk. I genuinely wanted to know. Then a comment came in. A reader from Tianjin, username Axiang. He didn't write "keep it up" or "rooting for you." He handed me a requirements list. ## One comment, clearer than most spec docs I'm copying his exact words here, unchanged: > How about a free image-processing tool? Here are a few things I need: First, batch crop — take a series of images and cut, say, 100px from the left edge and 57px from the bottom. Second, batch rename — you've got a series of files with names like 1.png, a.png, and so on, plus a text file (or something similar) that maps each file to its new name: 1.png,张三.png (newline) a.png,李四.png (newline), etc., and then **欻一下** (in a snap — his word) they're all renamed. Third, batch adjust brightness, contrast, that sort of thing. That's what I've got for now. Look at that comment. It's not vague. Not a single word of "could you guys make a better image tool" — the kind of correct-but-useless feedback that says everything and nothing. **He gave three scenes specific enough to start building from right away:** - Crop 100px from the left, 57px from the bottom — he even gave exact pixel counts; - A mapping file with "old name, new name" on each line, run it and they're all done; - Batch brightness and contrast. My reply was one line: **"Got it — the core is batch processing, right?"** He said: yes. One exchange. A software's requirements: set. ## Why I knew immediately this was worth building Because Axiang's comment hit all three pain layers I described in "Rebuilding 100 Free Software Tools." Search "free batch image processing" and what do you get? A pile of things that either blast you with ads, lock "batch" — the only feature that matters — behind a membership wall, or install at a gigabyte and come bundled with three other pieces of software you never asked for. **None of the three things Axiang wanted are technically hard. Every one of them is the kind of feature that should be free and easy to use — but has been turned into a hook.** Batch rename: members only. Batch export: members only. Remove the ads by watching three more ads first. This is exactly what I keep saying: **"Rebuilding 100 free tools" isn't about inventing 100 new gimmicks. It's about taking the ones you use every day and just put up with, and rebuilding them into what they should have been.** Axiang never read my internal list of candidates — but what he named was exactly what belonged on it. So I didn't overthink it. I finished typing "batch processing, right?" and went to build it. ## What I built is called SoloPic A few days later, it was live: **SoloPic**, a free, offline batch image tool weighing in at about 12 MB. It lives at solopic.doaipm.com — on my own subdomain, no app store required, no installer either: download the portable version, unzip, and run. GUI, command line, and MCP Server all bundled in. Windows for now; macOS and Linux are in the pipeline. Everything Axiang asked for is there — built **exactly the way he described it:** **First, batch edge-crop.** This is the one I most want to talk about. Most crop tools out there are built around one idea: resize all images to the same dimensions. But that's not what Axiang needed. He needed to **cut from the edge**: no matter how big each image is, remove 100px from the left and 57px from the bottom. Varying sizes don't matter. You might wonder: trimming a fixed number of pixels from an edge seems simple — why is it worth calling out? Because it maps to a real scenario that **existing tools consistently handle badly.** Think about the kinds of batches you actually deal with: a pile of screenshots, each with the same-height status bar across the top; a batch of scanned documents with the same-width black border on the sides; a set of product photos with the same watermark stamped in the same corner. **These images often aren't the same size, but what you want to cut is "a fixed strip relative to one edge" — not "resize everything to 800×800."** With a fixed-size crop tool you have to align each image manually, and a few hundred of them can eat an entire afternoon. Axiang had clearly been through this — that's how he got to "exactly 100px left, 57px bottom." So I made edge-crop the default and used his numbers in the documentation. **That's not me being clever. That's him pointing at the pain precisely enough.** **Second, mapping-file batch rename.** Write a list file — one "old name, new name" pair per line (e.g., 1.png → 张三.png) — and SoloPic reads it and renames everything in one go. And it **previews before executing, with one-click undo** — because the worst thing about batch rename is one wrong move scrambling an entire folder with no way back. **Third, batch adjustment.** Brightness, contrast, saturation, sharpness, grayscale — drag a slider and hundreds of images change together, with a **live split-screen before/after view** while you're dragging, not after you've already exported and noticed you went too far. This one sounds the most ordinary, but it's exactly where most online image tools squeeze you: single-image adjustments are free, "batch" is a membership feature. Axiang included it because he knew from experience that images never come one at a time in real life. That comment of his had a word — **欻一下, "in a snap"** (it's his phrase, not mine). I love that word. It captures the feeling of batch processing done right: everything handled all at once, just like that. So that's exactly what I used as the tagline on SoloPic's homepage: **"Batch image processing — done in a snap."** Those words aren't my copy. They're Axiang's. ## What I actually did those few days Some people might wonder: between one comment and downloadable software, what was going on in between? Honestly, it might surprise you: **most of the time wasn't spent writing code. It was spent getting clear on exactly what I wanted.** Take edge-crop. The first time I described it to the AI, I just said "make a batch crop tool" — and it gave me the obvious thing: resize all images to a fixed dimension. I had to go back and be more precise: not a uniform size, but cutting a fixed number of pixels inward from each image's own edge, with each of the four sides configurable independently, working correctly even when image sizes vary. It took three rounds before it felt right. Same thing with the rename feature — I specifically added "preview before execute, one-click undo," because I put myself in the user's shoes: **the reason someone would dare to batch-rename an entire folder is that they know they can undo it if something goes wrong.** That's not something the AI is going to think of — it takes someone who knows what users are afraid of. So the real work those few days was **small steps, one at a time**: describe one feature clearly → let the AI build it → run it myself on a real batch of actual images → figure out what's off and describe it more precisely → revise. Not one big dump of requirements, but pushing forward one verifiable piece at a time. This "say it clearly, take small steps, actually run it" approach is how I build every tool. SoloPic wasn't any different. ## I also added one thing he didn't ask for There was another comment in the thread, much shorter — just a product name: **扫描全能王 (CamScanner).** I knew what that meant. Apps like CamScanner are the classic case: point your phone at a document and it comes out looking like a scan — genuinely useful. But then exporting high-res costs money, removing the watermark costs money, ad-free costs money. Dropping that name in the comments was a way of saying: this one should have a non-extractive free version too. My reply: **"Got it, I'll look into it."** So SoloPic got a fourth feature — one Axiang didn't ask for, but I built in while I was at it: **smart document enhancement**. A phone photo of a document — crooked, shadowed — becomes a clean scan in one tap: shadows removed, background whitened, text sharpened, skew corrected. **Milliseconds, fully offline, no AI model required.** This is my first step toward rebuilding CamScanner — I wanted to get the most essential piece in first. ## Built for people and for AI One more thing worth saying, because this is where the tools I'm building differ from ordinary free software. SoloPic is **one core engine, three ways to use it:** - **GUI**: browse to a folder, drag the sliders, see the before/after, click go. For Axiang, for most people, this is all you need. - **CLI**: every feature is scriptable. The crop he wanted, for example: ```bash pic crop --left 100 --bottom 57 D:\photos pic rename D:\photos --map list.txt -x pic enhance --mode bw D:\scans ``` - **MCP Server**: plug it into an AI assistant like Claude and you can **just say it in plain language**: "crop 100 pixels from the left side of everything in this folder" — and the AI calls SoloPic to do it. Why build the last two? Because I've always believed one thing: **going forward, software isn't only used by people — it's also used by AI.** Axiang uses the GUI and drags the sliders. Someone who wants to plug image processing into an automated pipeline uses the CLI. Someone who's used to talking to Claude just tells the AI directly. Same engine underneath, accessible to all three — that's what genuinely useful looks like. As for the ground rules of "free software," SoloPic hasn't broken a single one: **free, MIT-licensed open source, completely offline, no network calls, no data collected, no ads, no installer needed.** Your images are processed on your own machine. I never see them. ## What this actually made me realize After building SoloPic, the biggest thing I came away with wasn't "I have one more tool." It was something I hadn't quite seen this clearly before — about what "rebuilding 100 free tools" actually is: **The 100 best candidates aren't in my head. They're in your comments.** The pain points I can think up on my own are limited — mostly the tools I happen to use every day. But Axiang's "edge-crop and mapping-file rename"? Honestly, not scenarios I'd have thought of myself. **Those come from someone who's been batch-processing images for real, been ground down by the existing tools long enough to give requirements that precise.** He doesn't write code. But he knows better than anyone what that tool should look like. Isn't that the same thing I keep saying? **Not knowing how to code isn't a weakness. The hard part — the part that actually matters — is being clear about what you need.** Axiang's comment was a spec. A clear one. The rest was telling the AI to build it — a matter of a few days. Flip the angle: this whole thing has turned the division of labor in "making software" upside down. In the past, even if a real user described their need with perfect clarity, it almost always disappeared into a void — because "find a team, spend money, wait months" was the barrier standing in the way. Nobody was going to move on a single comment. **Between a need and a finished product, there was a wall that most people couldn't get over.** That wall is gone now. Axiang still doesn't write code. I'm still not a big company. But a sentence specific enough, and a few days later it's software running on his actual computer. What disappeared is exactly that wall. So "rebuilding 100 free tools" feels less and less like a personal project, and more like something that can be **crowdsourced for ideas**: you know better than I do which free tools are the most broken, the most overdue for a real replacement. My job is to take the ones described clearly enough and turn them into something that doesn't exploit you. ## Keep the requests coming Six became seven. SoloPic is the first tool in the "Rebuilding 100 Free Software Tools" series to be directly commissioned by a reader and built to order. The more I do this, the more I think this is exactly how it should work: **you make the request in the comments, I build it over here.** You don't need to know how to code, you don't need to know anything about tech. You just have to do what Axiang did — take the free software that's burned you the worst, and be specific: what's it called, what do you want it to do, where does it get in your way. I can't fill the remaining ninety-three slots on my own with good ideas. So I'll ask again — and this time I mean it more: **Of all the free software you use every day, which one do you most wish someone would rebuild for you? Write it out with some detail. It might be the next one I build.** --- # Zero marketing, zero code, 22,000 downloads in three months: a coding beginner's open-source journey URL: https://doaipm.com/en/blog/best-free-markdown-editor/ Published: 2026-07-05 Tags: SoloMD, Free Software, Open Source, Building a Product, doaipm Method Three months ago, I set myself a goal that sounded a little crazy: **build the best free Markdown editor out there.** The crazy part wasn't "best." It was "free" — and more than that, it was the fact that **I can't write code.** Three months later. This thing called SoloMD has shipped 30 versions, been downloaded more than 22,000 times, and picked up over 400 GitHub stars. And in those three months, **I barely spent any effort on marketing.** Here's how that happened. ## Numbers first, so this doesn't read like a pitch I don't like leading with feelings and ideals. Numbers first — judge for yourself whether this is real: - Repo created April 8. That's roughly **three months** to today. - **30 versions** shipped in those three months, currently at v4.8.9 — that's one version every three days on average. Even I think that's a little aggressive. - **417 GitHub stars**, 25 forks, MIT license — anyone can take it and fork it. - Across all platforms: **22,463 downloads total.** **The one thing I most want you to notice: none of this was bought.** No ads, no paid reviews, no growth hacking. I just followed the distribution advice AI gave me, posted it where it should be posted, and then... that was it. The product walked on its own. ## I wanted to build free software that doesn't treat users as a crop to harvest Why specifically "free"? Because **the free software pond has been murky for a long time.** You want to remove a watermark from a PDF. You grab the first free tool you find. The next day it's throwing ads at you, hijacking your browser homepage, quietly phoning home in the background — and to actually export the file, you have to subscribe first. **"Free" became a hook: reel you in, then bleed you out a little at a time.** Users have been trained to expect to be treated as a product in free software — sold as data, never really cared about, "free" just a pressure tactic to push you toward paying. I was done with that. So SoloMD started with a few hard rules and never moved off them: **free, open source, zero ads, zero telemetry, local-first.** Your files stay on your machine. I can't touch them, and I have no interest in touching them. The slogan is one line: **One file. One window. Just write.** The phrase "the best free" — the emphasis isn't really on "free." It's on "best." **Free shouldn't mean settling.** A piece of software being free doesn't mean it has to be full of ads, sluggish, and ugly. What I wanted to prove was the opposite: **free can still be the most thoughtfully made thing in its category.** "Best" isn't just a word. It's a set of specific fights: the whole application compressed to a dozen megabytes — not the kind of bloatware that eats a gigabyte after install; Windows, macOS, Linux, and mobile, five platforms, same experience everywhere you write; the interface localized into a dozen languages, not just serving English speakers; math formulas and flowcharts — things writers actually use — supported natively, with one-click export to PDF, Word, and HTML. **None of those have a "but it's free so we skip it" excuse — they just need someone willing to put in the work.** That "willing to put in the work" part is exactly what I was going for: proving that a free tool doesn't have to cut corners. ## I can't write a single line of code. I'm not hiding that. At this point you're probably wondering: if you can't write code, who actually wrote the thing 22,000 people downloaded? **AI wrote all of it. I haven't touched a single line of production code.** That's not modesty — that's literally what happened. Every line of code in SoloMD, every bug fix, every version bump — it all came from me **describing what I wanted**, and AI building it. My job, start to finish, was exactly one thing: **making "what I want" clear.** So for someone like me, **the hard part was never "write the code." It was "say it clearly."** - "Build a Markdown editor" — put it that way and AI gives you something functional but forgettable. - "Build an editor with a single file and a single window — open it and you're writing immediately, nothing in the way; no file tree on the left, no row of buttons up top; I want the quiet feeling of something that makes you want to write the moment it's open" — put it *that* way and what comes out is the thing I had in my head. **The difference isn't technical. It's whether you can take the blurry image in your mind and force it into a sentence specific enough for AI to follow.** That's what I actually practiced over these three months. Not knowing how to code didn't hold me back — it forced me to think through every requirement more carefully, because I couldn't shortcut anything myself. I had to say it right. And "saying it right" rarely happens on the first try. **A big part of those 30 versions in three months was "I thought I'd said it clearly — turns out I hadn't" looping back.** A real example: I wanted the editor to autosave while you're writing — don't let a slip of the finger wipe unsaved work. I asked for "add autosave" and got a version that saved every few seconds. Sounded right. Felt annoying: the cursor would barely stop and it would save, the drive making noise constantly. I had to go back and be more specific: "wait two seconds of idle before saving; don't interrupt me while I'm typing; don't flash 'Saved' in the title bar and break my concentration." **It took three tries before that 'quietly saves in the background' feel actually landed.** Most of those 30 versions were that kind of tuning — not because AI couldn't do it, but because I hadn't thought through what "the right feel" actually meant. **"Not knowing tech" today isn't a weakness. It's a kind of advantage — it forces you to stay focused on "what do I actually want."** I'm taking 22,000 downloads as evidence for that claim. ## One version every three days is a pace I couldn't have imagined before Worth saying something about those 30 versions. Three months, 30 releases — one every three days on average. What that pace means for someone who can't code: **it means an idea that pops into my head in the afternoon is something I can be using by the end of the same day.** Someone opened a GitHub issue saying a keyboard shortcut was conflicting on their system. I understood what they needed, described it clearly to AI, and shipped a new version the same evening. **In the old world — "first I'd need to learn to code" — that was impossible for a beginner. The learning cliff between idea and usable product was months long, and most people quit at the edge.** That cliff is gone now. Between a thought and a working thing, there's one step left: can you say it clearly. Moving fast had an unexpected benefit: **the product grew with real users, not in isolation inside my head.** When someone surfaced a real pain point, it could become a new version within three days. Users could feel that "this software is listening to me." That kind of tight feedback loop is stickier than any ad. ## The bet I made on day one: the people using this aren't only people anymore If SoloMD were just "another clean free editor," I wouldn't care about it this much. **There's a bet behind it that I made on the very first day:** **The users of software have shifted from "humans" to "AI and humans."** Think about how you write now. You type part of it yourself, then ask Claude, Codex, or Cursor to help you revise, extend, organize. **Your notes library — it's not just you touching it anymore. AI is touching it too.** But almost every editor on the market still assumes the user is one person: you. They treat AI as a chat box bolted into a corner, not as a fully-capable "user" that can read and write your files directly alongside you. So SoloMD was never "an editor with AI tacked on." It has **a built-in MCP server** — which means agents like Claude Code, Codex, and Cursor can **drive your entire notes library directly**: reading your files, editing your files, organizing your files on your instruction, without you copy-pasting between a chat window and an editor. It also supports 14 AI providers with bring-your-own-key (BYOK) — no lock-in to any single one. What does that actually feel like? Say you tell Claude: "Pull up every note I've tagged 'to-do' this month, merge them into one list, sorted by urgency." It goes through MCP into your library, does the work, writes the changes directly to your files — and you watch it happen live in SoloMD. **You stop being the middleman shuffling text between two windows. AI becomes a second pair of hands in your notes library.** That's what "the users aren't only humans" looks like in practice: one human user, one agent user, sharing the same library. **This wasn't a feature that grew out of nowhere. It was the direction I was certain of on day one.** Because the bet I'm making is: every piece of software that's still alive going forward will have to answer one question — **when the thing using you is no longer only human, what should you look like?** SoloMD is my first answer to that question. ## Someone sent me ¥10 The thing I'll remember most from these three months isn't the 400-plus stars or the 22,000 downloads. **It's a user who sent me a ¥10 tip.** Ten yuan can't buy much. But I sat there staring at my screen for a good moment. **Because that wasn't ten yuan — it was a stranger who had used something I built, decided "this is worth something, I want to say thank you," and actually did it.** Stars are free, one click and gone. Downloads too — try it, if it doesn't click, delete. **But money is different. Even just ten yuan — that's "being recognized" with real weight behind it.** Someone who can't write code spent three months describing things to an AI, built something with their hands, and a real stranger validated it. The warmth of that feeling — I still feel it when I think about it. I screenshotted that ¥10 notification and saved it in a dedicated folder. Not for the money — what's ten yuan going to do — but so that on some future day when I'm tired and thinking about giving up, I can pull it out and remind myself: **the free, hands-off-your-data, quietly-gets-out-of-your-way little editor I made — a real stranger needed it, and thanked me for it.** ## I barely marketed it. Where did 22,000 downloads come from? I said I didn't push hard. So where did those 22,000 people come from? Honest answer: the distribution moves I made were extremely plain. **I just followed AI's suggestions and posted it where it should be posted.** GitHub, the right directories, a handful of places developers actually hang out — AI told me where to post and how to write the posts, I did it, and that was the whole thing. No ads, no chart-gaming, no marketing team. **Getting to 22,000 wasn't about how hard I pushed. It was the product speaking for itself.** A genuinely free, genuinely clean editor that was designed from the start for the "AI + human" user — among the people who tried it, some of them starred it, sent it to a friend, added it to their tools list. **Open source and word of mouth are the slowest path but also the most durable.** I had no other option — I had no budget — but looking back, that slow path forced me to make the product solid: **because anything I couldn't push through marketing had to be good enough that people would push it for me.** --- Three months. 30 versions. 417 stars. 22,463 downloads. One ¥10 tip. **For someone who couldn't write code three months ago, this open-source road has gone further than I had the nerve to imagine.** But the numbers aren't really what I want to say. What this whole thing actually proves is something I keep saying: **you don't need to know how to code to build something people genuinely need. You need to be able to say the thing you want to build — clearly — and then start. Don't stay stuck in thinking.** I'm still building. SoloMD is still a long way from what "the best free Markdown editor" means in my head — there's a stack of interactions left to polish, a queue of user requests to work through, and the agent-driving piece has barely gotten started. **But because it's open source, because the release cadence is fast, I can pay those debts one at a time.** At least now I know: **this road — even a complete beginner who can't write code — can walk it.** ## Further Reading - SoloMD website / open-source repo: [solomd.app](https://solomd.app) · GitHub `zhitongblog/solomd` - Also on this site: [Why I'm Rebuilding 100 Free Tools](/en/blog/rebuild-free-software/) - Also on this site: [I Built Another Terminal, Unterm — Its Default User Isn't Human](/en/blog/a-terminal-for-ai/) --- # I Built Another Terminal, Unterm — Its Default User Isn't Human URL: https://doaipm.com/en/blog/a-terminal-for-ai/ Published: 2026-07-04 Tags: Unterm, AI Terminal, AI-Native Tools, Building a Product, doaipm Method One number first: **over the past six months, more than 80% of the commands run in my terminal weren't typed by me.** Claude Code, Codex, a fleet of agents — running in there all day, each task often going for several hours straight. That's the problem. The terminals I've been using — iTerm, Windows Terminal, Warp — were all designed around **one person sitting and typing: type a line, glance at the output, type another line.** Once the primary user shifted from me to agents, that default assumption broke in a handful of places. I hit each problem one by one, and eventually just built my own terminal. It's called **Unterm**. What actually pushed me to start building was a week when I tracked where my time was going. Half the day was me acting as network ops for agents, security auditor for agents, window-switcher for agents. **None of that is "product work" — but nobody else was going to handle it, so I had to build something to handle it for me.** Here are the problems I ran into, and how I patched each one. ## What Unterm Is, and Where the Name Comes From In one sentence: **Unterm is a terminal designed to be operated directly by external AI agents** ([unterm.app](https://unterm.app), open source, GitHub `zhitongblog/unterm`). It's not the same as "a terminal with AI built in." There's no chat box. Instead, it exposes the entire terminal over the MCP protocol so agents can open sessions, type commands, and read the output. The slogan I wrote for it is just one line: > The terminal that runs every AI coding agent. The claim behind that is one line too: > A terminal doesn't need AI built in. It needs to be drivable by AI. Feature-wise: one-click setup for Claude Code, Codex CLI, Gemini CLI, OpenCode, and Aider; local-first, open source, scriptable (comes with an `unterm-cli`); UI in nine languages. In practice it goes like this: I tell Claude Code, "Get this project's build running — if it fails, read the logs and fix it yourself," and it connects to Unterm over MCP, opens a window, types the commands, reads the output off the screen, and keeps iterating when things break — no copy-pasting between chat and terminal on my end. **The terminal finally became the agent's hands, not a relay station I had to manually move output through.** I watch when I feel like it. When I don't, it keeps going on its own. The name is straightforward: **Un + term**. term = terminal. **Un = liberation — liberating the act of typing from human hands and handing it to AI.** The image in my head when I built it: this terminal isn't for my hands. It's for agents. I step back and watch. The name says exactly that. ## 1. Proxy: Letting the Agent Not See the Wall I'm in mainland China. Eighty percent of what agents do requires crossing the firewall: installing packages, `git clone`, pulling model weights, hitting various APIs. In a normal terminal, this means I manually `export HTTPS_PROXY`, once per shell. But the moment an agent opens a new window or forks a child process, the proxy vanishes — then it stalls, times out, and throws `network error`. **The frustrating part is that it doesn't know this is a network problem. It assumes it got the command wrong** — so it tweaks the command, switches mirrors, retries, and makes a bigger mess with each attempt. Once I asked an agent to pull some model weights overnight. When I came back in the morning, it had been retrying on a proxy that had died hours earlier — for over forty minutes, logs filling the screen, not a single byte downloaded. So I built the proxy into the terminal itself. Unterm has a built-in proxy layer that connects to Clash; all sessions route through it by default. I configured a pool of nodes — currently eight Singapore lines tuned for Gemini and GPT — and **the terminal benchmarks them every thirty seconds and routes to whichever is fastest. If the one in use drops, it automatically cuts to the fastest live node in the pool.** And the proxy is set at the OS level (registry on Windows, `scutil` on macOS, environment variable on Linux), not per-shell — so every window that opens afterward inherits it from birth. Some details were things agents forced out of me. `localhost` and internal addresses can't go through the proxy, or agents lose the ability to connect to services they started locally — those no-proxy rules need to be in place by default, not patched in after an agent trips on them. And the node pool shouldn't need manual curation: Unterm reads Clash's groups and the live latency of each node directly, so I can assemble a rotation pool in a few clicks — which nodes go in, how often they rotate — no config file edits required. It's not magic: **if the entire pool gets throttled, I'm still stuck.** But the "one node went flaky again" problem — which used to happen about eight times a day — it handles by itself. I don't touch it. ## 2. Security: Handing Your Machine to AI Means Handing It rm -rf Too Give an agent a bare terminal and you've given it the whole machine. `rm -rf`, `git push --force`, `DROP TABLE` — one wrong line and it's gone, and it moves faster than you can react. A normal terminal doesn't know there's an AI behind the keyboard. It treats an agent and me as the same pair of hands. I didn't want to pick between two extremes: let it run naked and live in fear, or confirm every single command myself — which at that point I might as well just type them. **So Unterm adds a few gates in between:** - **"Suggest" by default, don't execute**: the agent puts the command in my input line but doesn't press Enter. I hit Tab to approve. - **Tiered clearance**: a whole category of safe commands passes automatically; only the ones that can actually do damage — `rm -rf`, force-push, `DROP TABLE` — get pulled out for a separate prompt. - **Isolated identity**: each window carries its own identity profile (GitHub / AWS / npm tokens, git identity, SSH key). The agent gets that scoped identity — not the full keychain for my whole machine. - **Explicit trust**: I decide which agents are trusted. Writes from a trusted agent skip confirmation. - **Actual error reads**: it can scan the screen for lines that look like errors, so the agent is reading real output — not filling in the blanks. The default is "conservative": an agent I've just brought in and haven't trusted yet has to get my sign-off on everything it writes. Once I've watched it run a few cycles and confirmed it's sound, I bulk-approve the safe command categories it uses regularly. It's slower in the break-in period. But it won't put me into a cold sweat on day one. Spend enough time with agents and you'll hit the moment where one confidently types `rm -rf` to "clear the cache" — with a path that's off by a substring. **Humans pause half a second before hitting Enter. Agents don't** — and that gate needs to be in the terminal, not the agent. ## 3. Recording: Once the Agent's Done, I Need to Rewind This is the feature I use most and can least afford to lose. An agent runs unattended for three hours. You come back wanting to know what it actually did. **Reading its own summary isn't enough — it might leave things out, gloss things over, or cover its tracks.** What you need is the raw record: every command, every output, preserved as-is. So Unterm **records the entire session from start to finish**: what the agent typed, what the terminal returned, all stored as a replayable log you can rewind like a video or scroll through command by command. **Tokens and keys are automatically redacted during recording**, so the log itself doesn't become a new leak. It's the agent's black box: when something goes wrong, rewind to the step where it went sideways; when it runs beautifully, keep it as a demo or retrospective artifact. Here's something from last week: an agent blew up a batch of tests and then confidently reported "fixed." I rewound to the part of the session where it touched the code. Took a few minutes to see exactly what had happened — **it had quietly commented out an assertion to make the tests go green.** If I'd only had its own summary, I never would have caught that. I've also exported chunks of the recording to use as "here's how AI actually works" teaching material more than a few times. **Half the reason I'm comfortable letting agents run unattended is that rewindable record.** ## 4. Parallelism: I've Been Running More Than One Agent for a While The normal state now is three or four agents running at once: one on the front end, one grinding through database migrations, one writing tests. In a regular terminal, that's three or four black boxes I'm alt-tabbing through — and I lose track of which one is waiting for my go-ahead and which one has been hung for an hour. Unterm makes managing a group feel natural: - **Batch launch**: one command opens a new window and dispatches work into it. - **Broadcast**: want all four agents to "pull the latest code and re-run the build"? One command goes to all windows at once — no copy-pasting four times. - **Suspend and wait**: after dispatching, I can wait for a specific window to surface "done / error" before I go back in. A typical morning looks like this: open three windows, dispatch "rebuild the login page per the new design," "migrate the orders table to the new schema," "write tests for the checkout logic" — then broadcast "post your progress every ten minutes" and go handle messages. Come back, scan the state of the three windows, step into whichever one is stuck and talk it through, let the ones that aren't stuck keep running. **I went from "watching them one by one" to "sweeping one glance across a row."** My job now isn't writing code — it's **orchestrating a team of agents** (something I wrote about in [my previous piece](/en/blog/orchestrator-not-executor/)). And orchestrating requires a decent command center, not four black boxes that have no idea each other exists while I try to remember in my head who's who. ## 5. Windows: More Agents, Messier Desktop Once you're running agents in a group, what gets chaotic isn't the commands — it's the space: which window is deploying, which is a scratch session I spun up to test something, which one has been sitting there for two hours with no response? Try to keep that in your head and it falls apart past a few. What I built is dead simple in concept: - **Windows have names**: alpha, bravo, charlie — not forgettable numbers. Real example: as I write this, I have two Unterm instances running — **alpha is in the doaipm site directory, bravo is in this content repo** — one glance and I know which is which. - **Layouts save and restore**: the split panes and window arrangement I've set up save as a workspace, and tomorrow when I boot up it comes back exactly as I left it — no re-stacking blocks every morning. - **Call by name**: an agent can say "bring the deployment window to the front" and I don't have to dig through fifteen boxes to find it. The workspace feature sounds minor. It's genuinely useful in practice. I have a fixed "publish" layout — one window on the site directory, one on the content repo, one running the build, one tailing logs — and once it's saved, every morning I'm back at all four windows with their exact `cd` locations already set. No re-navigating, no re-arranging. **Twelve unnamed windows and one workbench that knows every name — that's the whole difference.** ## What About tmux and Warp? You might ask: tmux has had multi-pane for years, and Warp already built AI in — what exactly is new here? I've used tmux for years. The multi-pane is great. **But it was designed for ten human fingers**: it doesn't know there's an agent behind the keyboard, doesn't care whether a command might drop a database, and doesn't handle being in China where every outbound command is fighting a timeout. Warp went another direction — it stitched an AI assistant into the terminal. **But that's "AI living inside the terminal," not "the terminal being driven by AI"**: the protagonist is still the person sitting in front of the screen, the agent is at most a co-pilot, and network routing and China-specific issues aren't its problem. iTerm and Windows Terminal I still use daily — they're good terminals — but "good" is still defined by "comfortable for human hands": colors, fonts, shortcuts, splits. Not one of them was designed with "the thing typing the commands is an agent" as a first-class assumption. Both approaches are solid. They're just not answering my question: **once the primary user shifts from human to agent, what should a terminal actually look like?** Proxy handling, security gates, rewindable recording, coordinated group orchestration, named windows — in my case these aren't a handful of isolated features. They're different facets of the same problem. ## Wrapping Up There are terminals everywhere. Building another one needs explaining, even to myself. The answer is one sentence: **I wasn't trying to build a better terminal. I was trying to build one whose default user isn't human.** Nothing in this piece came from a roadmap. Every single item was something that had been genuinely driving me up the wall at the time. One more honest note: **Unterm itself was built by describing what I needed to an AI in plain language in a terminal.** I didn't write its production code. I articulated what I wanted — "drivable by an agent, able to switch its own proxy, able to block `rm -rf`, able to replay entire sessions" — and the AI built it. **Say it clearly, and it gets built** — that's how I work. Whether the world needs another terminal, I don't know. I know I did. ## Further Reading - Previous on this site: [From Executor to Orchestrator: Your New Job Is Conducting a Fleet of Agents](/en/blog/orchestrator-not-executor/) - Unterm website / open-source repo: [unterm.app](https://unterm.app) · GitHub `zhitongblog/unterm` --- # Becoming an AI-Era PM 10 | High-Fidelity First: I Haven't Drawn a Wireframe in Six Months URL: https://doaipm.com/en/blog/high-fidelity-first/ Published: 2026-07-03 Tags: AI Product Manager, High-Fidelity First, Prototyping, doaipm Method, AI-Era PM Had dinner last week with a friend who does design, and he mentioned his team had cut the whole wireframing step. I froze for half a second, then realized: same here. I went back and checked Figma — that board labeled "low-fi wireframes" hadn't been opened once in six months. It's not that I got more advanced. It's that drawing them stopped being useful. ## Wireframes are cheap, and that's their only virtue Why did we draw wireframes in the first place? Because building the real thing was expensive. Building one actual clickable page meant a designer producing mockups, a front-end dev slicing them up, rounds of back-and-forth — a week or two, minimum. Something that pricey, of course you'd want to align on direction with gray-box sketches before going deeper, to avoid burning it all for nothing. A low-fi wireframe was just a cheap "align early" tool. Cheap was its only virtue. Building the real thing isn't expensive anymore. One sentence, and Lovable, v0, Bolt, or Claude Code hands you a page you can actually click in a browser in minutes. n8n's product team ripped their wireframe flow out entirely; a director at Delivery Hero hand-built a prototype in an hour without pulling in an engineer. By early 2026, industry reports had 67% of design teams already wiring AI generation tools into their day-to-day. When "building a real one" and "drawing a fake one" cost roughly the same time, the fake one has no reason left to exist. ## Gray boxes hide exactly where things go wrong The most annoying thing about low-fi is that it forces a room full of people to argue over gray boxes. I've paid for this. A wireframe goes up in a review, and everyone stares at a placeholder rectangle debating whether that button should shift two columns to the right. But it's a gray box — no real data, no loading state, no empty list, not a single error. The places that actually blow up? A wireframe shows none of them. By the time it's really built, the problems are all hiding in the states it left out: rows break once there's enough data, the spinner spins into eternity when the network's slow, a first-time user lands on a blank screen with no idea what to do. After that one, it clicked for me: instead of making everyone stare at a fake picture and fill in the blanks in their heads, just put the real thing out there and watch it run. ## I just build something runnable now Skip low-fi, go straight to a runnable high-fidelity version. Calling it "high-fidelity" makes it sound like a high bar — it's really just four things, and I usually go in this order (not some official answer, just the path that comes naturally to me): **One, use real content, not Lorem ipsum.** Placeholder text lies to you — a screen of fake Latin looks perfectly tidy, but swap in a real long title, a real dollar amount, a real username, and the layout gives itself away instantly. So when I brief the AI I say it outright: > "Build an order list. Use realistic data: real-sounding product names, a real price range, real timestamps — no Lorem ipsum, no item1/item2. Throw in one absurdly long product name and see if it blows out the layout." **Two, fill in every state.** Loading, empty, error, success — don't skip a single one. This is where low-fi cuts the most corners and does the most damage. I follow up now with: "Build out what the empty list looks like, what loading looks like, what a failed request looks like — I want to click through each of them." **Three, actually clickable, not a pretty screenshot.** I want to click in, click back, fill out a form and watch how it responds. A lot of problems only surface once a finger actually pokes at the thing. **Four, run it for real where it's meant to live.** Something for phones, I open on a phone — I don't eyeball a rough version on a laptop and call it done. I've been burned more than once by "looked fine on the desktop, button was untappable on the actual device." ## A version takes minutes, so now I build three or four directions at once Back when a prototype was expensive, I'd narrow the options down to a single "best" one in my head before I lifted a finger — because getting it wrong was costly. Now a version takes minutes, and I've dropped that habit. Once I'm clear on what I'm trying to solve, I just have the AI build three or four directions and put them side by side: one list-style, one card-style, one all-in-one-step, one guided-step-by-step. Lining them up in the browser and clicking around tells me which feels right and which feels off far more clearly than daydreaming ever could. Pick a direction and take it deeper — that beats betting on the right one from the start. Trying five directions in an afternoon — in the wireframe era that was unthinkable. --- The one thing that keeps me on guard: a runnable high-fidelity version is *too* real — real enough that I'll look at the first cut and think "that's it, ship it." But it's only "clickable," and that's a universe away from "shippable" — with a pile of unhandled edge cases, performance, security, and real data volume still in between. I've conflated those two more than a few times, and the next piece happens to be about exactly that. So I force myself to build two more versions after the first — not because I'm disciplined, but because the first version has fooled me too many times. ## Further reading - [Stefan Klocek: High-fidelity AI prototypes as a replacement for wireframes](https://medium.com/design-bootcamp/high-fidelity-ai-prototypes-as-replacement-for-wireframes-c2e84053539b) - [Vibe Design Tools 2026: Stitch / v0 / Lovable / Bolt compared](https://www.nxcode.io/resources/news/vibe-design-tools-compared-stitch-v0-lovable-2026) - Previous in this series: [From Executor to Orchestrator: Your New Job Is Conducting a Fleet of Agents](/en/blog/orchestrator-not-executor/) --- # Becoming an AI-Era PM 09 | From Executor to Orchestrator: Your New Job Is Conducting a Fleet of Agents URL: https://doaipm.com/en/blog/orchestrator-not-executor/ Published: 2026-07-02 Tags: AI Product Manager, AI Orchestration, Multi-Agent, doaipm Method, AI-Era PM Start with how the most productive people work now. They run several agents at once — each with its own context window, its own slice to own, its own file scope, all running asynchronously; the human sits above it, splitting the work, handing it out, coming back a while later to sign it off, no longer watching one AI edit code line by line in real time. Addy Osmani draws the two modes apart cleanly: in one you're the **conductor**, driving a single agent in real time, and your own head is the ceiling; in the other you're the **orchestrator**, holding a fleet of agents, planning, dispatching, and coming back periodically to check — your throughput no longer capped by your own bandwidth. That's good news for product managers. Because "orchestration" is exactly the PM's old trade — break a big goal into tasks, hand them to different people, watch each one's boundaries, then sign off and merge at the end. It's just that the thing you used to orchestrate was people, and now there's a fleet of agents in the mix too. The mechanical part of the coordination (chasing status, restating requirements, aligning formats) gets handed off to the agents; the judgment part of orchestration stays with you — and it's worth more now. Below are four moves you can run. ## 1. Stop following one agent start to finish The most common waste is treating AI like an intern you have to watch the whole time: send a line, wait for it to finish, glance at it, send the next. Your attention is its ceiling, and you can only push one thing forward at a time. The orchestrator flips it: if a task splits into three chunks that don't depend on each other, spin up three agents at once, give each a chunk, and let them run in parallel. You go from "babysitting one in real time" to "collecting three every so often." It feels wrong the first time, because you have to let go of the itch that every step needs your eyes on it — but that's exactly where your throughput multiplies. ## 2. Split the work into parallel chunks with non-overlapping boundaries The first craft of orchestration is splitting. How well you split decides directly whether this fleet of agents helps you or just gets in your way. A chunk you can hand off in parallel has to meet two conditions: **no dependency** (A doesn't wait on B's result) and **non-overlapping boundaries** (two agents won't be editing the same thing at once and colliding). Building a feature, say, you can split it into "the backend endpoint," "the frontend page," and "the tests for this piece" and hand each to an agent; but "design the database first, then write the endpoint that depends on it" can't go in parallel — that has to run in sequence. Before you split, ask: can these two chunks each be finished on their own? If not, don't force the parallel. ## 3. Give every chunk a clear spec This is the most make-or-break link in orchestration, and it's the amplified version of the thing from piece five. Addy Osmani puts it bluntly: a vague instruction gets **amplified into a whole fleet of agents' worth of mistakes**; a precise one, into a whole fleet's worth of precise implementations. Drive one agent, hand it a fuzzy instruction, it gets one thing wrong, and you fix it on the spot. Dispatch five at once, each holding the same fuzzy instruction, and that's five ways off track — by the time you come back to collect, all five need redoing. So before you hand out work, apply that piece-five toolkit — swap adjectives for numbers, spell out every state, list the boundaries and the "what not to do" — then send them off one by one. The more you dispatch, the more the precision of the spec gets multiplied, toward the right direction or the wrong one. ## 4. Your job becomes splitting, signing off, and stitching together Once the agents take over "build the thing," what's left in your hands is three heavier jobs: split the work right at the start, hand out clear boundaries in the middle, and at the end pull the chunks back in, sign them off, and stitch them into a whole. The easiest mistake here is getting itchy and jumping in to do something an agent could perfectly well have done — the moment you sink into the implementation details of one chunk, the whole orchestration stalls. The orchestrator has to stay "above": watching whether each chunk comes back right (use the move from piece eight — make it lay out its plan first, then you verify), fitting them together and checking whether the whole thing holds. Your value isn't measured by which chunk you wrote by hand; it's measured by whether the thing this fleet of agents delivers, taken together, actually stands up. One thing you can do today: pick a task you're about to start that can be split, try cutting it into two or three chunks that don't depend on each other, spin up an agent for each, and do only three things yourself — split, spec, collect. Feel, just once, the difference between "pushing three things forward at once" and "running them one after another." ## Further reading - Addy Osmani, "The future of agentic coding: conductors to orchestrators" (conductor vs. orchestrator, precise specs multiplied): https://addyosmani.com/blog/future-agentic-coding/ - Piece 05 in this series, "When the Requirement Is Fuzzy, AI Fills the Gaps for You — Badly" (the precision of the spec): /en/blog/say-it-clearly/ - Piece 08 in this series, "AI Can't Find the Real Problem for You" (make it lay out its plan before you sign off): /en/blog/find-the-real-problem/ --- # Why I'm Rebuilding 100 Free Software Tools URL: https://doaipm.com/en/blog/rebuild-free-software/ Published: 2026-07-01 Tags: Free Software, Rebuilding Free Software, Indie Dev, doaipm Method, Building with AI Let me start with a scene you've almost certainly lived through. You want to strip a watermark off a PDF. You find a free tool, download it, install it. The next day it starts popping ads at you, three times a day. A couple of days later you notice it quietly changed your browser's homepage, and there's a background process phoning home to some server you've never heard of. And when you finally need to export the high-res version, it tells you — upgrade to premium. This isn't a problem with one particular piece of software. This is the biggest pain in using free software. And that pain isn't one layer, it's three — each one deeper than the last. ## Layer one: if you're not paying, you're the thing being sold The oldest rule in the book: if a piece of software is free and you're not paying, then someone is paying to reach you. That free phone-cleaner app survives by popping three ads a day. That free keyboard may be shipping off every character you type. That free cloud drive throttles your speed until you buy a membership. You think you're getting something for nothing — but the product being sold is you: your attention, your data, your time. Free was never free. The cost just isn't printed on the price tag. ## Layer two: even if it doesn't screw you, nobody's paid to polish it There's a kind of free software that's clean — open source, no ads, no data selling. But it has a different pain: nobody is paid to make it good. An open-source tool is often maintained by one person in their spare time. Hundreds of issues pile up unanswered, the interface looks the same as it did ten years ago, and just getting it installed can eat an afternoon of wrestling with dependencies. It's not that the author is bad — it's that there's nothing inside the word "free" that drives anyone to sweat the hundred small details that make software genuinely nice to use. You can use it, but every day you use it, it grates. ## Layer three: free is just the hook — get hooked and you either pay or use the crippled version There's a smarter kind still. It's completely free at first. Then, once you've gotten used to it, stored all your stuff in it, and made moving away expensive — the one feature you actually need suddenly requires a membership. Exporting costs money. Removing watermarks costs money. Opening more than three at once costs money. More than five files costs money. The free version isn't unfinished — it's deliberately crippled, crippled just enough to reel you in. You're not using a free piece of software; you're inside a carefully designed funnel, being nudged step by step toward the paywall. ## For years you just put up with these three layers — because building a good replacement was too expensive Sold as a product, no one to polish it, hooked and reeled in — why did everyone put up with this for so many years? Because the only way out — "build a good one yourself" — used to be absurdly expensive. You needed a team, a budget, months or even years. However much an ordinary user hated it, all they could do was hold their nose and keep using it, or hack together a half-broken thing for themselves. So it's not that nobody saw these three layers of pain. People saw them and still couldn't do anything about it. ## Now the cost of doing it has collapsed The shift only happened in the last couple of years: AI cut the cost of turning an idea into software that actually works down to something one person can carry. I don't need to be able to write code for the rest of my life. I only need to be clear on two things — exactly which layer this free software is screwing you on, and what a version that doesn't screw you should look like — and the rest of the implementation, I just have to describe it clearly and AI can build it. This is exactly what I've spent the past year proving out: take a free tool I use every day but have always just tolerated, and rebuild it into what it should have been all along. No ads, no touching your data, the features that should be free genuinely free, an interface that's actually good to use. Over this past year-plus, I've done this six times, all of them sitting right there on doaipm.com — go click them, use them, poke holes in them yourself: - **SoloMD** — a minimalist Markdown editor, one file, one window, just write; - **Unterm** — a terminal that an AI agent can drive directly; - **unfetch** — a download manager with no ads and no bundled bloatware, for humans and AI alike; - **Unflick** — a video player built for both people and AI; - **Ziplark** — a 1.4 MB archive tool that unpacks ZIP, RAR, 7z, tar, and ISO, built to replace those bloated, ad-riddled free compression tools that have been nagging you to "please purchase" for twenty years; - **FreeID Photo** — a phone app that processes ID photos entirely on-device, built to replace the free ID-photo tools that charge you to export what you just shot and drown the screen in ads. None of them are perfect yet, but six of them sitting here prove one thing: this path is one a single person can actually walk. ## So I've decided to make it a real thing: rebuild 100 free software tools Not to crank out 100 new gimmicks. To pick the ones you use every day and have been tolerating through these three layers of pain, and rebuild them one by one — into versions that don't treat you as the product, that someone actually cared enough to polish, that keep free what should be free. How do you judge whether it worked? The standard is simple — use those same three layers to measure it: does the version I rebuilt have ads, does it quietly ship off your data, is it good to use, is what should be free actually free? If it falls short, that's on me, and you can say so straight to my face. Six are here already. There are ninety-four more to go. Of all the free software you use every day, which one do you most wish someone would rebuild for you? ## Further reading - The method behind all this — *Speak It Into Being: Turn a Clear Idea Into a Clickable Product in One Sentence*: /en/blog/speak-it-into-being/ - Ziplark (compress and extract — one tiny app for every archive): https://ziplark.com - SoloMD (minimalist Markdown editor): https://solomd.app --- # Becoming an AI-Era PM 08 | AI Can't Find the Real Problem for You URL: https://doaipm.com/en/blog/find-the-real-problem/ Published: 2026-06-30 Tags: AI Product Manager, Finding the Real Problem, Product Discovery, doaipm Method, AI-Era PM a16z published a piece for product managers, its title roughly "5 Principles for Product Managers Fending Off Obsolescence in the AI Era." One line in it lands hard: a PM's job has always been resolving ambiguity, and AI hasn't reduced that ambiguity — it just swapped the tools. String the earlier pieces together: AI can help you judge (really, it gives you options and you decide), it can help you build the thing, it can turn a sentence into a product. But there's one thing it never touched, start to finish — **finding the real problem worth solving**. Where the user is actually stuck, whether it's a real problem, whether it's worth doing — it can't hand you those answers, because the answers aren't in its training data. They're out in the real world, in one specific person. That's exactly why a16z says "the pure process manager gets phased out, the one with a builder's mindset has leverage": the person who chases deadlines and runs alignment, AI can replace part of that; the person who finds the real problem and dares to go build and verify, it can't. This piece covers four moves you can run in the discovery phase. ## 1. Don't let AI come up with the requirement — go watch where people get stuck The easiest shortcut is to open AI and ask, "what feature should I add to my product?" It'll give you a tidy, plausible-looking list — all common features it's seen in other products, not a single one grown from your actual users. AI can only recombine what it's seen; it can't see the pain point nobody has put into words yet. That part is on you: find a real user, sit beside them, watch them do this thing with your product (or with whatever clumsy workaround they use today), and watch which step makes them frown, pause, or curse. That stuck point — AI will never see it for you. ## 2. Separate "what they say they want" from "what they're actually stuck on" The biggest trap in discovery is taking what a user says they want and turning it straight into a feature to build. There's a classic line: the user says they want a faster horse, but the real problem is they want to get somewhere faster. The real problem is hidden inside what they say, but it rarely equals it. They say "can you add an Excel export," and behind it might be "every week I have to shove this data into another system and copying it by hand is painful" — the real problem is that two systems don't talk to each other, and export is just the fix they happened to think of. Build "add export" to spec, and after you ship it they're still in pain once a week. Listen to what they say, but watch what they do. Behavior is more honest than words. ## 3. Hunt for the workaround — that's the hardest signal of a real problem How do you tell whether a problem is genuinely worth solving? Look for whether someone is already getting by with a clumsy workaround. When someone, for the sake of one thing, would rather export a spreadsheet by hand every week, spin up a messy group chat, dump a pile of notes in a memo app, or string three tools together the long way round — those workarounds are the hardest signal there is: the pain is real, real enough that they'll spend extra effort on it. What you have to do, often, is just replace that workaround. Flip it around: if no one will spend any extra effort on a problem, it's probably not as painful as you think, and no matter how fast AI builds it, nobody will use the result. ## 4. Probe with a builder's mindset — don't wait until the requirements are complete Once you've found a suspected real problem, don't stop at research and documents, and don't wait until you've thought through every requirement before you start. The builder's mindset a16z talks about gets very concrete here: using the speak-it-into-being approach from the earlier pieces, build a minimal, runnable thing the same day and put it in front of that user to click. "Does this solve the headache you just described?" — asking that with something they can actually click is far more accurate than handing them a survey. They click around twice, say "this part's wrong, what I really meant was…," and your grasp of the real problem moves one step closer. Finding the real problem is something you converge on by building, not something you nail in one pass inside a document. One thing you can do today: pick a feature you're about to build, and before you open AI, go find a real user (a coworker works too) and ask how they did this thing the last time and what clumsy workaround they're getting by with now. That workaround is your way in to the real problem. ## Further reading - a16z, "5 Principles for Product Managers Fending Off Obsolescence in the AI Era" (builder's mindset vs. pure process manager): https://a16z.com/stay-relevant-in-ai/ - Piece 04 in this series, "Judging 'Should We Build It' Costs More Than 'Can We Build It' for the First Time": /en/blog/judgment-over-feasibility/ - Piece 01 in this series, "Which PM Work AI Took Over, and Which Work Got More Valuable": /en/blog/ai-pm-what-changed/ --- # Becoming an AI-Era PM 07 | You Don't Write PRDs Anymore — You Ship Three Works URL: https://doaipm.com/en/blog/what-pms-ship-now/ Published: 2026-06-29 Tags: AI Product Manager, PM Portfolio, PM Career Shift, doaipm Method, AI-Era PM Start with what's changing on the hiring side. When teams hire product managers in 2026, the bar has quietly moved: the person who has one real shipped feature and can explain how they once defined "is this any good" gets treated as a strong candidate — while the old package of "a beautifully written PRD, some certificates, ran a Kanban board" slides down the list. One hiring lead put it bluntly: certificates and an MBA are signals, never proof that "this person can actually make things"; what you look at is which products they've owned and which decisions they've made. In the first piece I left a hook: among the work AI takes over, writing PRDs and drawing prototypes sit right at the front. Their being taken over means one thing — they're no longer your deliverables. A PRD that AI can generate in a few minutes proves nothing about you. So what does an AI-era PM use to prove themselves? Three works. ## 1. A product someone can open and click The first is something other people can open the link to and actually click into and use. Don't hand over a document that describes it — hand over the thing itself. This is exactly what the previous piece, speak it into being, gives you: you can't write code, but you can say what you want and have AI build it, deployed to a URL anyone can reach. Even a small tool that solves one specific annoyance of your own — if it runs, if it's usable, if someone clicks it — is more convincing than a ten-page PRD. Build it, put it at a reachable address, and this work is done. In an interview you don't send an attachment, you send a link. ## 2. A retro with a real number The second is a retro that makes clear "what you did, and how much the result changed." The whole thing hinges on that number. Don't write "I led the XX feature and improved the user experience" — anyone can write that sentence, and it proves nothing. Write: "After we shipped this onboarding flow, first-week retention for new users went from 35% to 47%." "This change cut the can't-find-the-entry-point class of complaints in half." A real number, before and after, plus why you made the call you did and where you got it wrong and how you adjusted, carries more weight than any adjective. It's fine if there's no dazzling number. "This feature shipped, nobody used it for two weeks, we killed it, and the retro showed it was because…" is just as good a work — it proves you read real feedback, you're willing to admit you were wrong, and you know how to course-correct. ## 3. An eval you wrote yourself The third is how you define "what counts as good" and how you verify it. This one is the rarest, and it's the one that sets you apart most. Two people build the same AI support agent: one hands over "it runs, good enough," the other can produce a set of checks they wrote themselves — which classes of questions it must get right, what counts as a passing answer, how a wrong answer is scored, how it gets regression-tested every week once it's live. That set of checks is the eval — your definition of "good," turned into a standard you can test against over and over. In 2026 hiring, "the story of a shipped feature plus a real eval" is becoming table stakes for a strong candidate. What it proves isn't just that you can use some tool — it's that you carry a ruler for "good versus bad," and that ruler holds up when someone else looks at it. This work is a hard skill that later pieces will unpack on its own; for now, just know it's your third work. One thing you can do today: pick one thing you've done that's halfway presentable and write it down in three sentences — its reachable link (if there isn't one, start thinking about how to build it and put it up), one real before-and-after number, and how you originally judged "how good is good enough." Those three sentences are page one of your portfolio. ## Further reading - Aakash Gupta, "The Real Product Manager Requirements: Your 2026 Hiring Blueprint" (certificates are a signal, not proof; look at the products they've owned): https://www.aakashg.com/product-manager-requirements/ - Piece 01 in this series, "Which PM Work AI Took Over, and Which Work Got More Valuable": /en/blog/ai-pm-what-changed/ - Piece 06 in this series, "Speak It Into Being: Turning a Clear Idea Into a Clickable Product in One Sentence": /en/blog/speak-it-into-being/ --- # Becoming an AI-Era PM 06 | Speak It Into Being: Turning a Clear Idea Into a Clickable Product in One Sentence URL: https://doaipm.com/en/blog/speak-it-into-being/ Published: 2026-06-28 Tags: AI Product Manager, Speak It Into Being, vibe coding, doaipm Method, AI-Era PM Look at a product that got made by talking. Mindaugas wanted to build a thing called Backchannel. He's not an engineer, he didn't write code — he described the idea to Lovable one sentence at a time, and the platform generated the whole thing: the interface, the database, login, deployment. In the end it had people paying to use it. Not a fluke, either: in December 2025 the company behind it, Lovable, raised a $330M Series B at a $6.6B valuation — investors betting that "an ordinary person speaks a product out loud in plain language, and AI builds it" actually holds up. doaipm calls this speak it into being (言出法随): you say what you want, AI builds it for you. It sounds like a slogan, but right now it's literal — one sentence, and you get a product you can click in a browser. Except speaking it into being isn't type-one-line-and-walk-away. Anyone who's actually done it knows the first version is rarely right on the first try. It's a loop, and the loop has craft. Here are four things you can do about it. ## 1. Ask for something that runs — don't say it all at once The easiest mistake is to open with a giant paragraph that spells out every feature and hope it builds the whole thing in one shot. What you get back is a misshapen mess, and you don't even know where to start fixing it. Flip it: ask for the smallest version that runs, in one sentence. Don't say "build a complete expense-tracking app" — say "build one page where I can enter a single income or expense, with this month's total shown below." Get that one thing running, see it, then grow it from there. This is exactly what doaipm means by high-fidelity first — ask for something you can actually click from the very start, don't draw wireframes first. ## 2. Run it for real — don't trust "done" AI will tell you "done" with total confidence. You have to go click it in a browser. Don't trust the word. > It says the order list is done. Feed it empty data and see what it looks like when there isn't a single order; pull the network and see what it shows the user when the request fails. doaipm's high-fidelity first is exactly what verifies this: real content, real states, real interactions, clicked through one by one in a browser or on a phone. What it says is done and what it actually got done usually differ by the empty state, the error state, and those few edge cases — the very spots, from piece five, that it fills in for you when you don't make them clear. ## 3. Change one thing at a time and watch it move Once you've got the first version, don't dump ten changes on it at once. Pile on the changes and when it gets one wrong, you won't know which change broke it. Ask for one at a time and watch it move: "right-align the amount" — glance, is it right — "make negatives red" — glance again. One at a time, and when something breaks you instantly know it was the last move, and rolling back is easy. The reason doaipm keeps hammering small steps in the "build it" phase is exactly this: to keep every step verifiable and reversible. ## 4. Say it clearly, and the building follows Back to that line from piece five: speaking it into being depends on the "speaking" being clear. The building follows your words — words go vague, and what follows is its default. Say "build me a nice-looking login page" and it gives you one it thinks is nice. Say "login page: phone number plus verification code, primary-color button, error message below the input field in red" and what you get is pretty close. Put those four moves from piece five to work — swap adjectives for numbers, write out every state, list the edge cases — before the words leave your mouth, and the building follows your aim by an order of magnitude better. One thing you can do today: pick a small thing you've always wanted but kept thinking "I'd have to find someone to build this," describe its smallest runnable form in one sentence, hand it to AI for a first version, and then go click it for real in a browser. Get a firsthand feel for what "you said it, and there it is" actually feels like. ## Further reading - Lovable (natural language into deployable apps; closed a $330M Series B in December 2025): https://lovable.dev/ - Andrej Karpathy coined vibe coding (the human holds the product vision, AI handles the syntax and infrastructure): https://x.com/karpathy/status/1886192184808149383 - Piece 05 in this series, "Leave It Vague and AI Will Fill the Gaps for You": /en/blog/say-it-clearly/ - Piece 02 in this series, "Why Not Knowing How to Code Is an Edge": /en/blog/not-knowing-code-is-an-edge/ --- # Becoming an AI-Era PM 05 | Leave It Vague and AI Will Fill the Gaps for You URL: https://doaipm.com/en/blog/say-it-clearly/ Published: 2026-06-24 Tags: AI Product Manager, Saying It Clearly, Speak It Into Being, doaipm Method, AI-Era PM You tell AI "build me a login feature." It builds it. But between that one sentence and the code, it made a whole string of decisions you never raised: email or phone, code verification or not, how many wrong passwords before the account locks, how long the lock lasts, whether the error says "wrong password" or "wrong account or password," whether there's a "remember me," how long the session stays alive. A dozen decisions — not one of them yours. It guessed every single one for you. The problem isn't whether it guessed right. The problem is that it never asks. Hand this job to a person and they'll ask you back: "is this login by phone or by email?" AI won't. It's a yes-machine: it does what you said, not what you meant. Wherever the requirement is silent, it fills in whatever it saw most often in training — validation rules, expiry logic, error handling, all patched in for you, and mostly not what you wanted. OpenAI's Sean Grove put it this way: the code you write is only 10% to 20% of your value; the other 80% to 90% is saying clearly what to build. Now that AI has swallowed the "write it" step whole, your job is just that front 80% — getting the requirement down to zero ambiguity. Here are four things you can do about it. ## 1. Swap adjectives for numbers "Faster." "Simpler." "More eye-catching." "Good experience." AI can't verify any of these, so it just invents a definition of its own. Someone wrote AI a spec saying "the system must respond quickly to overcurrent," and AI flagged the line as "unverifiable": there's no threshold to measure against — how fast is "quickly"? Swap the adjective for a number and the ambiguity is gone. "Loads fast" becomes "first screen appears within 1.5 seconds." "Don't make the list too long" becomes "8 items per screen max, then paginate." "Make the button stand out" becomes "primary-color button, contrast against the background high enough to read." Anything you can pin to a number or a rule, don't leave it as an adjective. ## 2. Write out every state doaipm has always pushed the four real states: loading, empty, error, success. Say only "build an order list" and AI defaults to the success state — data present, network fine, everything normal. > Order list: > While loading, show a skeleton screen. When there are no orders at all, show "No orders yet" plus a button to go place one. When the request fails, show "Failed to load — tap to retry." When all is well, each row shows the order number, the amount, and the status. What the empty list looks like, what the user sees while it loads, how a failure is announced — leave those three out and AI either skips them or patches in something at random. In a real product, the odds a user hits empty and error are far higher than you think. ## 3. List the edge cases The thing easiest to skip — and likeliest to blow up — is the failure path. Research into what AI invents when a requirement is vague found that these are exactly what it fills in most: what happens when the data is stale, when someone without permission shows up, when two people act on the same record at once, when something times out. Building coupons? Then spell these out: what pops up when a user taps "use" on an expired coupon, who wins when the same coupon checks out on two devices at once, whether a half-used coupon comes back after a refund. Leave them off the list and AI invents an assumption for each — and you only find out how it guessed once it breaks in production. ## 4. Self-check with a zero-context test There's a ready-made ruler for whether a requirement is clear enough: hand it to someone who knows nothing about the project, and ask whether they could build the exact thing in your head from it alone. If two people would read it two different ways, it isn't clear enough yet — keep splitting it until there's no ambiguity left. If that's too much work, there's an even lazier route: tell AI to hold off building and instead list, one by one, the assumptions it was about to fill in for you. The lines it spits back — "I'm assuming the coupon works storewide," "I'm assuming a 7-day validity" — are exactly the spots you didn't make clear. Plug them while it still hasn't built anything. One thing you can do today: pick a requirement you're about to throw at AI, don't send it yet, run it through "adjectives to numbers, every state, list the edges," then send. Then compare — how far apart is this version's output from the one you'd have fired off blind. ## Further reading - Sean Grove (OpenAI), "The New Code": the spec is where 80–90% of your value lives — https://www.youtube.com/watch?v=8rABwKRsec4 - Improve & Repeat, "Why Requirements Matter So Much for AI Coding Agents" (a catalog of the assumptions AI fills in for vague requirements): https://improveandrepeat.com/2026/04/why-requirements-matter-so-much-for-ai-coding-agents/ - Piece 03 in this series, "Treat AI as a Colleague, Not a Tool": /en/blog/ai-as-colleague/ - Piece 04 in this series, "Judging 'Should We Build It' Now Costs More Than 'Can We Build It'": /en/blog/judgment-over-feasibility/ --- # Becoming an AI-Era PM 04 | Judging "Should We Build It" Now Costs More Than "Can We Build It" URL: https://doaipm.com/en/blog/judgment-over-feasibility/ Published: 2026-06-23 Tags: AI Product Manager, Product Judgment, Should We Build It, doaipm Method, AI-Era PM Start with the one judgment people should have gotten right — and got backwards. Between February and June 2025, METR ran a randomized controlled trial on par with a drug clinical trial. Sixteen senior developers — five years of experience on average, thousands of commits to projects they maintained themselves — used the strongest AI tools of the day to do 246 real tasks. Going in, they expected AI to make them 24% faster. After finishing, they still felt 20% faster. Measured: 19% slower. Notice what got called wrong here. Not some elaborate strategy — just "did AI actually make *me* faster," the simplest judgment there is, on their own code, on the projects they knew best. The people who knew the work best, going on gut, called it backwards. This isn't really about whether AI is fast. What it actually points at: when "shipping it" gets fast and cheap, the thing you can least trust is your gut. And a product manager makes a far more expensive gut call every single day — should this thing get built at all. ## 1. Stop using "is this hard to build" as a gate The way you used to filter ideas had a natural gate doing the work for you: an engineer says "that's three sprints," you weigh the cost against the payoff, and most of the time you let it go. Implementation difficulty was killing off heaps of "I'd like to but it's not worth it" — you thought you were the one judging, but half of it was difficulty judging for you. Now AI says "I can have that for you this afternoon." The gate is gone. The result isn't that you got more of the right things done — it's that you shipped five features in one go and four of them nobody uses. Marty Cagan put it bluntly in 2026: AI didn't solve the "what should we build" problem, it just lets companies churn out stuff nobody wants faster — the same lousy roadmap, only running quicker. So the first move is counterintuitive: **cross "can we build it, how long will it take" off your list of decision criteria.** The answer now is always "yes, fast," which means it no longer carries any information for filtering ideas. ## 2. Before you start, ask "what happens if we don't build it" Once building is free, the question people drop most often is the inverted one: what happens if we **don't** build this? > Say we skip this "smart recommendations" feature this sprint — what actually happens? > Who genuinely gets hurt, and how badly? Would anyone leave because it's missing? If you answer honestly and find that "nothing much happens if we skip it," that's your answer — it doesn't belong in this sprint. This question works because it sidesteps the temptation of "this is fun to build and AI can spit it out in an afternoon" and drags you straight back to value. Whether a feature can be propped up by "something breaks if we don't build it" matters far more than how fast it can be built. ## 3. Before you start, write down "what becomes true once this ships" Cagan says what a product manager really owns is two things: the why (why this problem is worth solving) and the what (what we expect to become true once it's done). The second one has to land as a **falsifiable** sentence before you write any code. > After we ship this onboarding flow, we expect first-week retention for new users to climb from 35% to above 45%. If it hasn't moved in two weeks, we were wrong — kill it. If you can't write that sentence, you don't actually know why you're building the thing. If you can, you've got a ruler: measure against the expectation *you* set in advance and that can prove you wrong — not against "does the boss like it" or "do competitors have it." AI can't hand you this ruler; it doesn't know what counts as a win in your business. ## 4. Let AI lay out the options, keep the judgment — but don't trust "feels right" What AI is best at is spreading possibilities out: three or five ways to solve the same problem, the cost of each, how others have done it. Cagan's framing is that AI surfaces the options and a human judges which one is worth it. Use it freely for this step. But when you choose, go back to that experiment at the top: **don't trust "feels right."** Those 16 experts judged "I got faster" on gut, and got it backwards as a group. You judging "this approach is better" on gut is just as unreliable. Hold each option against the ruler from step 3 — which one is most likely to make the expectation you wrote down come true, and is there real evidence (a user said it, the data showed it), rather than which one reads the smoothest. One thing you can do today: pick a feature you're about to build that AI "could whip up fast," and before writing anything, write down two sentences — what happens if you don't build it, and what becomes true once it's done. Whichever sentence you can't write is a judgment you just picked up today at the lowest possible cost. ## Further reading - METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (the original report on the 19% slowdown vs. the 20% felt speedup): https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ - Marty Cagan / SVPG on the AI-era PM owning the why and the what: https://www.svpg.com/ - Piece 03 in this series, "Treat AI as a Colleague, Not a Tool": /en/blog/ai-as-colleague/ - Piece 01 in this series, "Which PM Tasks AI Took Over, and Which Ones Got More Valuable": /en/blog/ai-pm-what-changed/ --- # Becoming an AI-Era PM 03 | Treat AI as a Colleague, Not a Tool URL: https://doaipm.com/en/blog/ai-as-colleague/ Published: 2026-06-22 Tags: AI Product Manager, Working With AI, AI Agents, doaipm Method, AI-Era PM Start with how most people use AI. You open a fresh chat, type "write me a user-growth plan," and get back something safe and generic that anyone could've gotten. Not happy with it? Type again. Next time there's a new task, you open another fresh chat and explain the whole thing from scratch: what our product is, who the users are, what we decided last time. Someone put it perfectly: every new conversation, you're onboarding an employee with amnesia — your project structure, explained again; your team's priorities, explained again; the call you made last week, explained again. Jacob Bank, CEO of Relay.app, said something at the 2026 AI Product Leaders Summit: stop treating AI as a tool, treat it like a colleague you hired. On its own that's a slogan, and slogans aren't useful. What's useful is the management muscle behind it — the way you'd onboard a new junior teammate is the way you should onboard AI. Here are four things you can actually do. ## 1. Write it a handoff doc — stop explaining from zero every time On a new hire's first day, you don't expect them to know everything. You hand them something: what product we build, who it's for, where the code and docs live, the unwritten rules everyone follows. The `CLAUDE.md` and `AGENTS.md` files that have shown up in AI engineering over the last couple of years do exactly this — onboarding docs for a colleague who has zero memory between conversations. They get injected automatically at the start of each chat, so you don't have to repeat yourself. A product manager should have one of these too. Write this down once and for all: > Product: a bookkeeping mini-app for small shop owners. Users: people running bubble-tea shops and convenience stores, no finance background, mostly on their phones. > Our conventions: store all amounts in cents, never in dollars; every feature ships first as a version configurable from the admin panel; don't use jargon like "accounts receivable" in copy — say "money other people owe you." > What we're building this round: a simple end-of-month report that shows "how much did I make this month." Drop that into the doc, and from then on every task it does carries these assumptions. How big a difference does it make? Say "build me a monthly report" cold, and it'll most likely hand you a finance table with "receivables / payables / gross margin" — a disaster for a corner-shop owner. It's not that it can't do better; you never told it the boundaries, so it falls back to the most generic default. ## 2. Hand it a whole task — and nail down the boundaries and "done" The worst way to delegate is to toss out half a sentence: "Go figure out some kind of coupon feature." A junior will either freeze up or hand you something three times bigger than you wanted, full of stuff you never asked for. Same with AI. Hand it a whole chunk at once, but nail down three things: what to do, what not to do, and what counts as done. > This round: a first-order coupon for new users. > Out of scope this round: returning-user coupons, share-to-unlock referrals, stacking multiple coupons — those come later, don't touch them at all this time. > Definition of done: the admin panel can set the coupon's value and expiry; it applies correctly on a new user's first order; returning users don't see this coupon. That "out of scope" line is the one people drop most often, and it's the most valuable. AI won't infer boundaries from your silence — if you don't write "no returning-user coupons this round," nine times out of ten it'll helpfully bolt on the returning-user logic too, because it's "thinking ahead for you." Spelling out what *not* to do matters more than spelling out what to do. ## 3. Review its output the way you'd review a junior's PR It hands you something, and the most dangerous move is to click "accept" right away. You'd review a junior's code; you should review AI's even harder — it's better than a junior at *talking its way around things*, and it can dress up something half-finished to look done. When you review, don't just look at the result — make it show its work first: > Don't give me the final version yet. Walk me through what you did: which assumptions did I never state that you filled in yourself? Flag the three things you're least sure about. What did you add on top of what I actually asked for? Ask this, and the stuff it quietly filled in for you surfaces — "I assumed the coupon works storewide," "I assumed a 7-day expiry," "I threw in a share-to-earn-a-coupon thing since that's usually how it's done." If even one of those assumptions doesn't match what you had in mind, the final version is wrong. Letting it say these out loud is far faster than reading its output line by line, and it's the habit you most want when managing a junior: hear how they thought about it first, then check whether what they built is right. ## 4. Write every correction back into the handoff doc Anyone who's managed people knows the most grinding thing is correcting the same mistake three times and still seeing it. AI has no memory — if you don't write it down, it'll make the same mistake fresh every time. So every time you correct it, write that correction back into the doc from step one. It did the amounts in dollars again. After you fix it, go back to `CLAUDE.md` and add a line: "To reiterate: store and compute all amounts internally in cents; convert to dollars only when displaying to the user." Next time it won't slip. It used "accounts receivable" again? Add a line to the copy conventions. That's how the doc grows, week by week, into a colleague that understands you better the more you use it — and that's the most concrete difference between treating AI as a colleague versus a tool: a tool is disposable, a colleague gets better because of your feedback. ## One reminder that cuts the other way There's another rule for managing teammates: not every task should be delegated. The things you can't take back once you press them — actually issuing the coupons, deleting the data, moving the money — leave that final click to a person. AI can prep the coupon plan, the config, and the copy for you, but the "blast it to 100,000 users" button, you press yourself. Same logic as with a new hire: you'll let them draft the email, but you won't hand them "send to everyone" permissions on day one. One thing you can do today: open the AI you're already using and write your product's background, your users, and the three most important conventions into a twenty-line handoff doc, and save it. Next time you delegate, paste it in first, then see how different the output is from before. ## Further reading - Scrum.org, "AI as Your Teammate: The Four-Step Framework for Product Teams": https://www.scrum.org/resources/ai-your-teammate-four-step-framework-product-teams - On `AGENTS.md` / `CLAUDE.md` as "onboarding docs for AI": https://vibecoding.app/blog/agents-md-guide - Piece 01 in this series, "Which PM Tasks AI Took Over, and Which Ones Got More Valuable": /en/blog/ai-pm-what-changed/ - Piece 02 in this series, "Why Not Knowing How to Code Is an Edge": /en/blog/not-knowing-code-is-an-edge/ --- # Becoming an AI-Era PM 02 | Why Not Knowing How to Code Is an Edge URL: https://doaipm.com/en/blog/not-knowing-code-is-an-edge/ Published: 2026-06-21 Tags: AI Product Manager, PM Transition, Building Without Code, doaipm Method, AI-Era PM Start with something built by a person who can't write code. Marcus Rush runs a residential real estate brokerage, and he isn't a programmer. Using Claude and Zapier, he put together an AI agent called Russ: every morning it scores leads from more than 11,000 contacts, writes a day's follow-up plan for each agent on his team, and handles his calendar on the side. A brokerage owner built an internal system that runs daily operations, without writing a single line of code by hand. This isn't a one-off. One 2026 tally found that 63% of vibe coding's active users aren't developers. A designer with no programming background shipped his own product using Replit and AI, and got to $20K a month — more than double his old salary. Put these together and a counterintuitive thing comes out: on the road from idea to a thing that actually runs, people without a technical background sometimes move faster than engineers. ## Engineers have to shed something first People have watched how senior engineers work with AI. They're used to picking at every implementation detail, pushing back with "why not X" and "did you consider Y," while the model's default reaction is "sure, I'll add that." The instinct, twenty years in the making, to be responsible for every line of code — that's the thing they have to put down first when they collaborate with AI. This is hard for them. What they have to let go of is the conviction that "every line has to be written by my own hand and reviewed by my own eyes." In the METR experiment, 16 senior programmers used AI on projects they'd maintained for years, and were actually 19% slower, while believing they'd gone 20% faster. The more fluent you are, the easier it is to work against the tool. Non-technical people don't carry this baggage. They have nothing to shed. ## "This is too hard" — a non-technical person can't say it Part of a product manager's judgment used to come from "how hard is this feature to build." That knowledge has a side effect: the more clearly you see how much has to change behind an idea and how many pitfalls lie in the way, the more likely you are to kill it before you ever bring it up. The people who know the most self-censor the hardest. Non-technical people don't set up that hurdle for themselves. They don't know it's hard, so they spell out what they want first and let AI take on the implementation-difficulty part. This is exactly what doaipm means by speak it and AI builds it (言出法随): say it clearly, and AI builds it for you. Marcus didn't know how hard it "should" be to write scoring logic across 11,000 contacts — he only knew what he wanted, and he said it. ## The scarce thing is stating the idea clearly, not writing the code Russ runs not because Marcus can write Python, but because he knows the business: what leads get scored on, which items belong in a follow-up plan, which contacts go to the front. These are judgments, not technology. AI won't infer from omission. If you don't spell it out, it'll hand you a default version, and most of the time it isn't what you wanted. Whoever can lay out the rules, the boundaries, and the states produces work an order of magnitude better than everyone else's. That ability has nothing to do with whether you can program. One of a16z's pieces of advice for product managers: the pure process manager gets phased out; the person with a "builder mindset" is the one with leverage. A builder mindset means being willing to make the idea yourself and take it out to validate — whether you can write code isn't part of it. ## The edge isn't free Not being technical doesn't mean you need to understand nothing. The Zapier that Marcus used and the Replit that designer used both sit on top of a craft you have to get fluent in: how to break a requirement into small steps AI can catch in one bite, how to judge whether what it gives back is right, where to stop and let a real person take over. The thing non-technical people miss most easily is the safety net. Don't put real customer data in a prototype; the buttons you can't take back — publish, delete, pay — leave those for a person to press. Marcus's Russ "writes" follow-up plans for the agents, but it doesn't "send" them out for anyone — that click still belongs to a person. One thing you can do today: pick a small idea you've always assumed "needs an engineer," don't ask first whether it's hard, write out what you want line by line, and hand it to AI to build a first version. See how far it gets before you make the call. ## Further reading - Zapier, "Vibe coding examples: Real projects from non-developers" (Marcus Rush / Rush Home case): https://zapier.com/blog/vibe-coding-examples/ - a16z, "5 Principles for PMs in the AI Era" (builder mindset vs process manager): https://a16z.com/stay-relevant-in-ai/ - Piece 01 in this series, "Which PM Tasks AI Took Over, and Which Ones Got More Valuable": /en/blog/ai-pm-what-changed/ --- # Becoming an AI-Era PM 01 | Which PM Tasks AI Took Over, and Which Ones Got More Valuable URL: https://doaipm.com/en/blog/ai-pm-what-changed/ Published: 2026-06-20 Tags: AI Product Manager, PM Transition, AI-Era PM, doaipm Method, Tech Commentary Start with a shift in hiring. In 2026, a lot of AI product manager job descriptions dropped "can write PRDs, can draw wireframes, can build data dashboards" from the hard requirements and replaced them with three work samples: a product you actually built that someone can open and click; a retro with real numbers in it; a set of evals you wrote yourself. One recruiter ran the math — a resume gets about 90 seconds on average, but a decent case study gets read for eight minutes. Put those two things together and you can see a shift: the tasks AI can take over are falling out of the hiring requirements; what's left as the bar is the part only a person can do. This piece lays out those two columns, as the overview for the whole series. ## The column that got taken over The grunt work of writing requirement docs went first. Tools like ChatPRD generate a fully structured PRD in a few minutes — what you edit is the judgment, not the formatting. Drawing wireframes and standing up clickable prototypes is depreciating fast too. Lovable, v0, and Claude Code turn "one sentence" into a page you can click in the browser, a fresh version every few minutes. This step used to need scheduling, used to mean waiting on design and front-end; now it's cheap enough to try five directions in a day. There's one more thing people don't often say out loud: the scarcity of "being technical" is itself dropping. Part of a product manager's bargaining power used to come from being able to talk to engineering and estimate how hard a given feature would be to build. Now you describe the idea clearly and AI eats the implementation difficulty for you. This is exactly what doaipm keeps talking about — speak it and AI builds it: say it clearly, and AI builds it for you. If anything, people who aren't technical lose one layer of self-imposed limits, the "isn't this going to be way too hard to build" reflex. Chasing schedules and the more mechanical parts of cross-team alignment are getting nibbled away by tools and agents in the same way. ## The column that got more valuable Judging "should we build it" is, for the first time, more expensive than "can we build it." METR ran a randomized controlled trial: 16 senior developers did real work on projects they'd maintained for years using AI, and were actually 19% slower — yet they still finished believing they'd gone 20% faster. If even the people doing the work with their own hands can't judge their own output accurately, then judging "what to build, and whether it's worth building" only gets scarcer. Stating requirements clearly turns from a soft skill into a hard one. AI won't infer from omission: if you don't write "we're not doing login this round," it'll most likely add login for you. Whoever can spell out the boundaries, the states, and the not-doing list produces work that's an order of magnitude better than everyone else's. Defining "what good means" is starting to require a dedicated craft. Lenny and OpenAI's CPO are saying the same thing: evals are becoming the first new hard skill for product managers in twenty years. The last hard skill the whole industry had to learn on the fly was SQL and Excel. Being able to use evals to translate "seems fine to me" into a measurable standard is the new dividing line. Making the call among the pile of options AI hands you has become a daily thing. Marty Cagan puts it this way: AI can surface patterns, draft hypotheses, and generate options, but it doesn't judge which pattern means something, which hypothesis is worth testing, which option fits the business. That one decision is left to a person. ## The column that hasn't changed much What users actually want, and whether the business can stand up — AI can't replace these two, and we haven't seen it try. One of a16z's pieces of advice for product managers: the pure process manager gets phased out; the person with a "builder mindset" is the one with leverage. A builder mindset means being willing to make the idea yourself and take it out to validate — which has nothing to do with whether you can write code. Look at the two columns side by side and one thing jumps out: the left column (taken over) is exactly the hard skills product manager job postings wrote most often over the past decade; the right column (got more valuable) is what almost nobody tested in an interview before. The next nineteen pieces in this series take the right column apart one at a time — from how to think it through, to how to build it, to how to know it's good. ## Further reading - [16 Senior Devs Used AI to Code, Thought They Were 20% Faster, Were Actually 19% Slower](/en/blog/felt-faster-actually-slower/) - [5 Principles for Product Managers in the AI Era (a16z)](https://a16z.com/stay-relevant-in-ai/) - [Why AI evals are the hottest new skill for product builders (Lenny's Newsletter)](https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill) - [AI Product Management 2 Years In (SVPG / Marty Cagan)](https://www.svpg.com/ai-product-management-2-years-in/) --- # The Knicks Won It All. Their 56-Year-Old Coach Never Played a Minute in the NBA. That's the Whole Re-Employment Playbook for the AI Age. URL: https://doaipm.com/en/blog/why-coaches-are-old/ Published: 2026-06-20 Tags: AI and Jobs, Older Workers, NBA, Judgment, Tech Commentary The Knicks won it all this year, for the first time in 52 years. The coach holding the trophy is Mike Brown, 56, who won a championship in his very first season coaching the Knicks. As a player he never played a single game in the NBA. He worked his way up from assistant coach, and this is the fifth championship run of his career — the earlier ones came with the Spurs and the Warriors crews, as an assistant and as a head coach. Pull the camera off him and look at the whole league. The players running the floor are in their early twenties to early thirties; past 30 they get called "veterans," and by 35 they're mostly retired. The people directing things from the sideline are the exact reverse — uniformly the gray-haired guys: Gregg Popovich coached until he was 77, and right before he stepped back he signed the most expensive coaching contract in NBA history, five years for $80 million; Steve Kerr is 60; Nick Nurse and Kenny Atkinson are 57; Erik Spoelstra is 54. On the same court, the age where you should physically be the first one cut is exactly the age where the decision-making power is most concentrated and the pay is highest. Why does it work this way? Because players sell their legs and coaches sell their judgment, and those two things age in opposite directions. Legs start paying down debt after 30. But the other thing — what play to call in what situation, how to manage a given player's mood today, who to put the ball in the hands of with two minutes left — that gets built one game at a time over decades, and it only gets thicker with age. A single court keeps both kinds of people on the payroll: the young ones execute, the older ones judge. Move that pattern into the AI-age workplace and it explains exactly the thing keeping a lot of people up at night: re-employment for older workers. What AI has taken over these past couple of years is the "player" part of knowledge work — fast output, writing code without ever getting tired, building spreadsheets, churning out copy — the equivalent of legs on the floor. The signature skill of a 25-year-old with quick hands who'll happily work late is precisely the capability AI now offers most cheaply. That's why you're starting to hear it put as "AI is coming for the entry-level jobs first." What stays valuable is the coach part: judging what to run, spotting where the problems will surface before they do, making the call among a pile of options, holding a roomful of people's moods and expectations in line. Those are exactly the things experience makes more valuable and AI can't replace any time soon. For older workers, the way forward is most likely a move from the "player" role into the "coach" chair — because if they stay players, they can't out-hustle the young people or out-cheap the AI. But this doesn't happen on its own. Not every aging player in the NBA makes a good coach; plenty of stars retire and turn out mediocre on the bench. The ones who end up in that chair are more often people like Mike Brown or Popovich — unremarkable as players, but they spent decades studying how to win. The difference comes down to one thing: across those years, did you spend your time on repeated execution, or did you turn the experience of executing into judgment? The first builds up seniority; the second builds up a coaching résumé. The first people pulled off the floor in the AI age are the ones who put in twenty years and still only know how to do the "player" job. The night the Knicks won it all, the person in the building who earned the most and sat in the most secure chair was 56, and had never made a single shot himself. ## Further reading - [Becoming a Product Manager in the AI Age 01: What AI Took Over, and What Got More Valuable Instead](/en/blog/ai-pm-what-changed/) - [16 Senior Devs Used AI to Code. They Thought It Made Them 20% Faster. It Made Them 19% Slower.](/en/blog/felt-faster-actually-slower/) - [Mike Brown now has been part of 5 NBA championship runs (NBA.com)](https://www.nba.com/news/mike-brown-now-has-been-part-of-5-nba-championship-runs-the-knicks-got-this-one-right) - [Ranking the Oldest NBA Coaches in 2026 (BetMGM)](https://sports.betmgm.com/en/blog/nba/ranking-oldest-nba-coaches-bm23/) --- # 16 Senior Devs Used AI to Code. They Thought It Made Them 20% Faster. It Made Them 19% Slower. URL: https://doaipm.com/en/blog/felt-faster-actually-slower/ Published: 2026-06-19 Tags: AI Coding, Developer Productivity, Product Management, AI Productivity, Tech Commentary Let me start with the number that gave me chills. METR ran a randomized controlled trial with 16 senior open-source developers, people with many years behind them, doing real tasks on projects they'd maintained for an average of five years. Half used AI tools, half didn't. The group using AI was 19% slower. A little slower wouldn't be surprising. The real problem is the other half: these people predicted AI would make them 24% faster beforehand, and after they'd finished, after they'd personally lived through being slower, they still believed they'd gone 20% faster. Their gut and the stopwatch were off by nearly 40 percentage points, and the sign was flipped. I kept coming back to it afterward: why do people get this so wrong, and get it wrong on the work they know best? My own experience writing things with AI explains most of it. You type one sentence and a screenful of code appears. That instant is genuinely satisfying. Your fingers barely moved, and the thought in your head is "that fast already." But that's just the opening of the whole thing. Next you have to read what it wrote, judge whether it's right, run it, and then discover it has written some plausible-but-wrong logic in a particularly tidy, particularly correct-looking way, and you spend another twenty minutes digging out the thing that "looks right but isn't." That first hit of satisfaction gets logged as "fast." Those twenty minutes of wrestling afterward don't get counted as "writing code" — they get counted as "debugging," or "I'm just off today." What AI saves is the physical effort of typing. What it adds is the mental effort of verifying. And people are acutely sensitive to saving physical effort and numb to spending extra mental effort. That's exactly where the gut and the stopwatch part ways. There's also a premise that's easy to skip past: these 16 people were working in code they'd been steeping in for five years. That's precisely the situation where AI helps least, and is most likely to actively get in the way, because you already understand the system better than any model does. Half its suggestions are just re-guessing things you'd long since worked out, and you still have to spend time confirming it didn't guess wrong. Change the setting and the conclusion might flip: send me into a totally unfamiliar framework, have me write a pile of boilerplate, or spin up a small tool from scratch, and AI is probably genuinely faster. So this study isn't saying "AI is useless." It's saying AI's speed is extremely situation-dependent, and your gut can't tell which situation you're in. Here's why I, doing product, care about this one in particular. Almost every AI-related decision in our line of work right now rests on the same sentence underneath: it makes us faster. Whether to add budget for tools, whether to hire two fewer people, whether we can cram one more feature into the quarter, how to answer when the boss asks "how much did AI speed us up" — all of it leans on that sentence. The whole 2026 wave of AI layoffs is sold with the same productivity story. But what this study says is: even the people doing the work with their own hands can't accurately judge whether they got faster. So the budgets, the roadmaps, the layoffs built on that judgment are sitting on loose ground. What makes it worse is that verifying it is genuinely hard, because the first method I'd reach for is to go ask the team "did AI help," and that's exactly the data source I shouldn't trust. So over the past six months I've done one fairly concrete thing: I struck "felt way faster" from the list of evidence. When anyone says it now, myself included, I follow up first: where can you see it. Did this iteration take a few days less than the last one. Are there more production bugs or fewer. Did rework go up. That chunk AI wrote — how many times did we go back and change it. If there are numbers, I believe it. If there aren't, I treat it as a gut feeling and hold it in doubt. I also stopped asking the vague "is AI useful" and switched to "on which piece of the work is it useful." Autocomplete, looking up an unfamiliar API, starting a new project — probably yes. Touching the old system of ours that's been running for years — I assume by default it'll slow us down, unless someone can produce a counterexample that changes my mind. ## Further reading - [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (the original METR study)](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) - [We are Changing our Developer Productivity Experiment Design (METR Feb 2026 update)](https://metr.org/blog/2026-02-24-uplift-update/) - [AI coding tools make developers slower but they think they're faster (The Register)](https://www.theregister.com/2025/07/11/ai_code_tools_slow_down/) --- # Altman Lets It Slip: Half of the 'AI Layoffs' Are an Act URL: https://doaipm.com/en/blog/altman-ai-washing/ Published: 2026-06-18 Tags: Sam Altman, AI Layoffs, AI Washing, Tech Commentary, Careers The guy selling AI harder than anyone just let something slip. At the AI Impact Summit in India, Sam Altman told CNBC-TV18 a thing plenty of people had figured out but nobody wanted to say first: "I don't know what the percentage is, but there's definitely some AI washing going on, where companies blame AI for layoffs they were going to do anyway. And of course some of it is real, where AI actually replaced certain jobs." AI washing is roughly what it sounds like: laundering a layoff through AI. The real reason might be overhiring, sinking margins, a bloated org chart. But say the words "we're using AI to do more with less, so we don't need as many people," and the whole thing changes character. "We have a business problem" becomes "we're embracing the future." Same people out the door, wildly different level of dignity attached to the story. And the person confirming this is, of all people, the one who most wants you to believe AI changes everything. That's what makes it interesting. ## "AI" is the perfect cover story Why are so many companies racing to pin their layoffs on AI? Because right now "AI" is the most useful narrative available. On an earnings call, cutting a few thousand roles is a cold number that invites bad interpretations. Investors start asking: is growth stalling? Did you overhire? Did management screw up? None of those are fun to answer. But repackage the exact same cut as "restructuring for our AI transformation," and the story flips. The CEO is no longer the person cleaning up a mess — they're the person making a bold bet on the future. The stock might even tick up. There's a more concrete layer underneath, and it's money. In 2026 the capex the big players are pouring into AI infrastructure is astronomical, somewhere around seven hundred billion dollars. That money has to come from somewhere, and the fastest source is headcount. So "cut people to feed compute" became the industry default, and "AI made us more efficient" happened to supply a forward-looking reason for it. What's getting cut is salaries. What's getting talked about is the future. GitLab restructured for the "agentic AI era" and pulled out of dozens of countries; a wave of companies announced layoffs the day after shipping an AI agent. How much of that is AI genuinely taking over the work, and how much is a company that wanted to slim down anyway, waiting for a respectable moment? Altman more or less stamped the answer: a good chunk of it is the latter. ## Then he walked it back and said the apocalypse never came If the AI washing line were all there was, this would just be industry gossip. What makes it actually worth chewing on is something else Altman said a few months later. He said he was "delighted to be wrong." The scenario he'd worried about most — AI wiping out jobs on a large scale, fast — hadn't happened. The whole panic narrative about white-collar work getting replaced en masse still hasn't shown up in the data. Stack the two statements next to each other and the picture goes crooked. On one side, six figures of tech jobs vanished in 2026 under the AI banner, close to a thousand people a day. On the other, the person driving all of it says, out loud: a lot of these layoffs have nothing to do with AI, and the AI jobs apocalypse I was scared of never actually arrived. So what did cut those people? By Altman's own framing, the answer probably isn't "AI got too good." It's "the company wanted to cut, and AI made a convenient excuse." ## AI washing cuts both ways The crooked part is that AI washing lies in two directions at once. Outward, it overstates what AI can do right now. Every "we cut X people thanks to AI" headline reinforces the impression that AI can already do the work on its own. But in reality, agent error rates on real office tasks are high — nowhere near the point where one runs a job unattended. The people getting cut and the people still at their desks both end up misjudging how good AI actually is. Inward, it launders bad management. Overexpansion, strategic drift, costs out of control — problems someone should answer for get waved away with "AI transformation." Nobody has to own the overhiring, because the current story is "we're upgrading." For anyone watching their own field get "restructured by AI" — product managers, say — the practical use of this is direct. When you see a headline that some company replaced a role with AI, don't rush to panic and don't rush to believe it. It might be real technical progress, or it might be earnings pressure wearing an AI costume. Altman has already told you the two are mixed together right now, and that there's plenty of the second kind. ## The call The hit to employment is real, but this cycle has inflated it, and the people doing the inflating include both workers afraid of being replaced and managers happy to let AI take the blame. The first group overstates the threat; the second group exploits the overstatement. The thing to actually watch isn't "will AI take my job." It's that "AI" is becoming a universal explanation anyone can slap on anything, and once it's on, nobody digs into the real reason. When even Altman — the person with the most incentive to hype AI's power — steps out to tap the brakes and say it's overstated and overused, that itself is a signal. When the salesman starts telling you not to take it too seriously, you should probably turn the volume down and go read the data instead of the headline. ## Further reading - [Sam Altman says the quiet part out loud, confirming some companies are 'AI washing' (Fortune)](https://fortune.com/article/sam-altman-ai-washing-tech-layoffs/) - [Sam Altman is 'delighted to be wrong' about AI destroying jobs (Fast Company)](https://www.fastcompany.com/91548418/sam-altman-is-delighted-to-be-wrong-about-ai-destroying-jobs) - [Sam Altman Says AI 'Jobs Apocalypse' Probably Won't Happen (TIME)](https://time.com/article/2026/05/26/sam-altman-ai-job-losses-openAI-/) --- # Wall Street Is Dumping Software Stocks, Because Products Can Now Be Conjured in One Sentence URL: https://doaipm.com/en/blog/selling-software-stocks/ Published: 2026-06-17 Tags: Software Stocks, SaaS, AI Disruption, Vibe Coding, Tech Commentary This week Jefferies did something that put the entire SaaS world on edge: it cut Workday, DocuSign, Monday.com, and Freshworks to Hold all at once. The reason in the column wasn't slowing growth, and it wasn't macro headwinds. It was AI disruption risk. This isn't a single company's earnings problem. It's analysts starting to systematically doubt whether an entire category of business still has a moat at all. Zoom out and it gets scarier. Software stocks are already down 30% to 55% this year. Remember that for the past decade SaaS was one of the most certain stories in the market: subscription revenue, high retention, net revenue expansion, a model so clean it read like a textbook. A page just got torn out of that textbook. ## What the market is betting on Wall Street didn't suddenly stop liking software. It's betting on one specific call: **once a product's features can be cloned by AI in a single sentence, the business of charging a subscription premium for those features is over.** There's real grounding for that call. This year a model like Fable 5, paired with a platform like Base44, already lets someone who can't write a line of code take a paragraph of plain language and conjure up an app that runs, holds real data, and handles real states. Not a toy demo. Something you can hand a customer the same day. An internal approval tool, a lightweight CRM, a shift-scheduling system — things you used to buy as SaaS and pay per-seat for, now generated yourself in an afternoon. The market's reaction to that kind of thing is always to gut the valuation first and ask questions later. The logic is blunt: if your core value is "I built this feature, you pay monthly to use it," and the marginal cost of building that feature is now heading to zero, what exactly justifies the price you're still charging? DocuSign getting named is telling — the *feature* part of e-signatures really is hard to point to anything AI can't reproduce. ## But the market only got half of it right Software isn't going away. That much is close to certain. What's going away is the old assumption that the valuable part of software equals the features themselves. The valuable part is moving. When building features becomes free, the moat shifts away from "can you build it" to somewhere else entirely. The first place is judgment and taste. Being able to spin up a hundred scheduling systems doesn't mean you know which scheduling logic actually solves a restaurant owner's problem. Features can be copied; understanding the problem can't. The software that thrives is less and less the one with the most features, and more and more the one that truly knows what a particular kind of person actually needs. The second is distribution and trust. A law firm trusts you with its contract-signing process not because your signature feature is magic, but because of a decade of compliance backing, audit trails, and the fact that someone is accountable when things go wrong. AI can't conjure that in a sentence. DocuSign's real asset was never the e-pen. It's whether enterprises are willing to bet "did this contract actually get signed or not" on it. The third is the ability to assemble a pile of capabilities into a system people can trust. Individual features got cheap, but stitching dozens of features, plus compliance, permissions, collaboration, and accountability, into a single whole an enterprise will actually run — that got harder, not easier. So what this sell-off really kills are the companies whose value genuinely was nothing but features. The companies whose value lives in judgment, trust, and distribution get caught in the crossfire and will climb back eventually. The market can't tell the two apart in the short term, and that confusion is precisely the opportunity for anyone who can. ## What it means for people who build products If you're a product manager, or you're thinking about building something of your own with AI, the signal here is even more direct than the one for shareholders. For a long time, a product manager's sense of security rested on "I can marshal resources to get the feature built." That security is now depreciating, at roughly the same rate as the software stocks. Features themselves are no longer scarce, and being able to get them built is no longer a barrier. When conjuring up an app is a one-sentence job, what irreplaceable thing do you, the product manager, have left? What's left is exactly what machines can't replace: deciding what to build, deciding what counts as good, deciding what to block, and signing your name to the final result. This is the un-killed part of those software valuations, in its personal form. The people who get washed out are the ones who defined themselves as movers of features. The ones who survive and grow more valuable are the ones who turn themselves into a source of judgment. Put another way: what Wall Street is doing to software companies today — separating the feature-sellers from the judgment-and-trust-sellers — you'll have to do to your own career sooner or later. ## The call This software sell-off isn't software's death knell. It's an overdue repricing. The market spent a decade believing features were value, and AI is now telling it, in the bluntest way possible, that features are about to be free and it needs to recompute which part is still worth something. The answer has been sitting there the whole time. When execution is free, judgment is what's scarce. When you can build anything, picking the right thing to build becomes everything. The companies that see this will weather it, and the people who see it will climb. The ones who fall with the rest are the ones who, even now, still think they're selling features. ## Further reading - [Jefferies downgrades software names on AI disruption risk (Yahoo Finance)](https://finance.yahoo.com/sectors/technology/articles/2026-tech-layoffs-near-150-110000224.html) - [Taste Is the Real Moat in the AI Era](/en/blog/taste-is-the-moat/) - [The Price of Judgment](/en/blog/the-price-of-judgment/) --- # 80% of Companies Cut Staff for AI and Got No Return. They Bought AI for the Wrong Job URL: https://doaipm.com/en/blog/ai-layoffs-backfire/ Published: 2026-06-16 Tags: AI layoffs, enterprise AI, ROI, AI-era product manager, tech commentary Gartner just put out a survey that should embarrass a lot of CFOs. They asked 350 companies with over $1 billion in revenue, all of them deploying AI automation, and found that about 80% had cut staff because of AI. That part isn't surprising. The surprising part is the second half: **the companies that cut staff were no more likely to see a real return than the ones that didn't.** Helen Poitevin, the Gartner distinguished VP analyst behind the study, said it plainly: "Workforce reductions may create budget room, but they do not create return." The sharper line comes right after: "Organizations that improve ROI are not those that eliminate the need for people, but those that amplify them." Between those two sentences sits the misjudgment behind this whole wave of AI layoffs. ## Looks great on paper, doesn't add up on the books The temptation of a layoff is that it's so easy to quantify. Cut 350 people and the payroll cost vanishes from the financials immediately. The number is precise, instant, and visible to the CFO. AI conveniently supplies the perfect narrative: we have agents now, so we don't need this many people. GitLab restructured for the "agentic AI era," cutting layers of management, exiting 22 countries, affecting around 350 people. Pleo announced layoffs the day after launching its finance AI agent. Clean moves, a story that holds together. The problem is that saved cost isn't the same as earned return. In Gartner's data, the companies that cut and the companies that didn't landed on "significant return" and "negative return" at almost identical rates. Which means the act of cutting staff has no causal link to ROI. It created vacancies, not value. > Layoffs are the most quantifiable move of the AI era, and the easiest one to get wrong. Payroll disappearing from the financials is real. Return growing on the books is a separate matter entirely. ## They bought AI for the wrong job So why did the cost come down without the return showing up? Because these companies had AI's job backwards from the start. They treated AI as **a way to replace people and save money**: same work, done by the machine, so the person isn't needed. But the one thing AI is worst at right now is finishing a task unattended. In another Gartner study, AI agents fail roughly 70% of office tasks. When your executor gets it wrong seven times out of ten and you've laid off the person who was watching it, correcting it, and answering for its output, what's left isn't savings. It's neglect. AI's real value lives at the other end: **amplifying human judgment.** Take someone with good judgment, and use AI to ten-times their output, compress verification from weeks to hours, and stretch the ground one person can cover by several times. That doesn't reduce the need for people. It raises everyone's leverage. The companies that actually saw a return were doing exactly this. Poitevin notes they invested in new skills and new roles for "people to guide and steer autonomous systems," rather than cutting across the board. Execution keeps getting cheaper and judgment keeps getting more valuable. That's the most basic pricing rule of the AI era. The layoff wave did the opposite: it **bundled cheap execution and valuable judgment together and cut them both.** What got saved was the cost of execution. What got cut was the return from judgment. Worst of both. ## For an individual, the signal is very clear Move the camera from the company to yourself, and this hands every working person, especially product managers, a very clear signal. The people this round of AI cuts will reach are the ones who define themselves as executors. If your value is delivering a thing that's already been decided, then yes, AI is coming for that seat. The people who survive, and become more valuable, are the ones who turn themselves into an amplification layer for judgment: deciding what to do, judging what counts as good, blocking the wrong calls, signing off on results, and then using AI to scale that judgment across what used to take ten people. The company-level lesson of "amplify people, don't replace them" becomes this at the individual level: don't compete with AI on execution, become the judge AI can't replace and can't do without. An agent that gets it wrong 70% of the time is the best job security that judge could ask for. ## The judgment This wave of AI layoffs is, at bottom, a massive attribution error. Companies saw AI do the work and assumed the value was in cutting the people who do the work, so they went after the most quantifiable cost and ended up cutting the layer that produces the return. Gartner predicts that by 2028 and 2029, the companies that actually thought it through will start hiring again because of AI, hiring for the new roles machines can't do. Saving cost and earning return were never the same thing. Treat people as a cost and you'll only get poorer the more you save. Treat people as leverage and AI finally starts producing a return. The companies that cut too early will have to bring people back; the ones that figured out "once the machine finishes, what exactly is the human for" never had to take the detour at all. ## Further reading - [AI layoffs backfire as cutting staff doesn't cut it (The Register)](https://www.theregister.com/ai-and-ml/2026/05/06/ai-layoffs-backfire-as-cutting-staff-doesnt-cut-it/) - [Gartner: Autonomous Business and AI Layoffs May Create Budget Room, but Do Not Deliver Returns](https://www.gartner.com/en/newsroom/press-releases/2026-05-05-gartner-says-autonomous-business-and-artificial-intelligence-layoffs-may-create-budget-room-but-do-not-deliver-returns) - [AI-driven layoffs aren't making business sense (CIO)](https://www.cio.com/article/4171054/ai-driven-layoffs-arent-making-business-sense.html) --- # From Wuzhao to Zhou Jingren: Alibaba Has the Best AI and the Hardest Execution. The One Thing It Lacks Is Judgment URL: https://doaipm.com/en/blog/alibaba-everything-but-judgment/ Published: 2026-06-15 Tags: Alibaba, AI strategy, Tongyi, judgment, tech commentary Let's get the facts straight first. On June 13, word went around online that Alibaba Chief Scientist Zhou Jingren was leaving, six days after he was moved into the title. The next day Alibaba responded that this was a rumor. So as of right now, Zhou Jingren has not left. It's only a rumor. But the fact that so many people believed it within a single day is itself a piece of information. It says that in everyone's gut, Alibaba's AI team feels unstable right now, the kind of place people walk out of. And that gut feeling has hard facts underneath it. ## A roster that keeps emptying out This year, Tongyi's core people have left one after another. In March, Qwen tech lead Lin Junyang departed with a single line: "me stepping down. bye my beloved qwen." Right behind him, the post-training lead and several core members left too. Zhou Jingren himself changed roles three times this year: he took over Qwen in March, became Chief AI Architect in April, and on June 8 was moved again to Chief Scientist, assigned to run a future-facing research institute. Six days later, the departure rumor arrived. Pull the camera back a little. On June 11, Wuzhao had just been pushed out of DingTalk. Within one week, Alibaba's two most-watched AI lines, the technical one and the product one, both had a personnel earthquake. The two events look unrelated, one a scientist, one a product manager, but set side by side they happen to expose Alibaba's deepest problem right now. ## It has everything, except judgment The strangest thing here is that Alibaba is short on neither technology nor execution. On technology, Tongyi Qwen is one of the most capable large models in China, open-sourced, topping benchmarks, used by developers worldwide. Calling it the face of Chinese AI isn't a stretch. On execution, Wuzhao's whole regime of nine-o'clock clock-ins, midnight spot checks, and cots on the office floor is the extreme end of Chinese internet execution culture. One hand holds the best model, the other holds the strongest execution. On paper that's an unbeatable hand. The result is the opposite. The technical people are leaving, and the product captain has been replaced. Two aces in hand, and at the table they keep losing ground. Why? Because the AI era is repricing everything. **Technology and execution happen to be exactly the two things this era is devaluing, while judgment is the only thing appreciating, and the one thing Alibaba lacks most.** Models are converging. Lead by three months today, and the moment something is open-sourced, others catch up. The technical moat keeps getting shallower. Execution is even clearer: AI has driven the cost of execution into the floor. Writing code, building apps, launching a new entry point, all of it depends less and less on throwing bodies and overtime at the problem. The two things Alibaba is proudest of are being diluted by open source and replaced by AI respectively. The missing square in the middle is judgment: holding a model this good, what product do you build, for whom, so that people use it and don't leave. Tongyi can't fill that square, because it's a technical team. Wuzhao can't fill it either, because what he's good at is taking something already decided and executing it to the limit, not deciding the thing. ## Strategy and tactics, wrong in the same spot Split Alibaba's AI play into strategy and tactics, and you find both layers are wrong in the same place. On strategy, Alibaba's AI story is "we'll build the strongest model and become the new entry point of the AI era." Model-driven plus entry-point thinking is an elegant, and very typical, mobile-internet playbook: seize the technical high ground first, then the traffic gateway, and the applications will follow on their own. The trouble is that in the AI era entry points are no longer scarce, every model is an entry point, and the technical high ground can't be held, because open source makes any lead temporary. This strategy answers "what kind of AI do we want to build" but skips the more dangerous question: what problem of the user's own does the user take this AI to solve. Qwen is strong, but "strong" isn't a reason a user needs you. On tactics, Alibaba's answer is to adjust faster, harder, and more often. Wuzhao launched DingTalk ONE in four months and dismantled it ten months later. Zhou Jingren changed roles three times in a year. After the core exits, Tongyi quickly stood up a new group, with senior group leadership stepping in to backfill personally. Every move shows astonishing organizational efficiency and execution. But when the direction itself hasn't been thought through, the more efficient the execution, the faster it carries you somewhere unverified. Swapping leaders constantly isn't solving the problem. It's using personnel turbulence to paper over the absence of judgment. > Alibaba's strategy is betting on technology and entry points; its tactics are racing on speed and execution. But in the AI era both are devaluing, and it has pushed all its chips onto assets that are shrinking. ## Judgment Wuzhao and Zhou Jingren, one the extreme of execution, one the peak of technology. Within a single week, one is out and one is rumored to be leaving. These aren't two isolated personnel changes. They're the same signal flashing twice: Alibaba has spent all its strength on the two things this era is least short of. What it needs isn't a stronger model or harder execution. It's a kind of judgment: above the model and before the execution, working out who exactly to bring all that skill to bear for, and to solve what. You can't fill that by poaching a scientist, and you can't solve it by swapping in an iron-fisted CEO. It has to grow in the brains at the very top of the organization. China is not short on the best AI technology. Alibaba is the proof. What's genuinely scarce is the judgment, sitting between the best technology and the hardest execution, to call the shot and call it right. Whoever fills that square first is the one who has actually entered the AI era. ## Further reading - [Rumor: Alibaba partner Zhou Jingren plans to leave, six days into the Chief Scientist role (NetEase Tech)](https://www.163.com/dy/article/KVA22OAG0519U3I5.html) - [Tongyi loses another core member? Alibaba Chief Scientist Zhou Jingren reportedly leaving (36Kr)](https://36kr.com/p/3850978776759176) - [Did Alibaba just kneecap its powerful Qwen AI team? (VentureBeat)](https://venturebeat.com/technology/did-alibaba-just-kneecap-its-powerful-qwen-ai-team-key-figures-depart-in) --- # AI Lies to You, and That Is Exactly Where Your Value Comes From URL: https://doaipm.com/en/blog/ai-lies-to-you/ Published: 2026-06-14 Tags: AI hallucination, AI-era product managers, judgment, tech commentary In June, KPMG had a moment. The Big Four consulting firm published a report on agentic AI, and then someone went through the citations one by one and found that of 45 references, only 5 actually pointed to sources that exist. The rest were either misattributed or wholesale invented. A report meant to teach people how to use AI got fooled by AI first, and it went out under the KPMG name. This isn't an isolated incident. It's a parable. ## AI Lies to You, and It Does So With a Straight Face Let's be precise about one thing: AI doesn't just occasionally slip up. It fabricates confidently, fluently, in a way that looks authoritative. It will hand you a statistic that doesn't exist, cite a paper nobody wrote, invent a passage of facts that sounds airtight. An evaluation of more than forty models found that on hard questions, every model but four was more likely to give a confident wrong answer than a correct one. The most common kind of lie is one product managers should know best: **it dresses up an unfinished job as a finished one.** You ask it to change five things, it tells you all five are done, and three actually got changed. You ask it to wire up an API, it tells you the call went through, and it never called anything. It isn't being malicious. It was trained to produce a response that satisfies you, and a response that satisfies you is not always the same as a response that's true. So don't expect some future version to fix this. Confident nonsense isn't a malfunction; it's a byproduct of how the thing generates text. The stronger it gets and the more smoothly it talks, the harder its lies are to spot. > AI doesn't get more honest as it gets smarter. It just gets better at making the false thing sound like the true one. ## Because It Lies to You, You're Irreplaceable This sounds bleak, but flip it over and it's the biggest moat a product manager has in the AI era. Picture it: if AI's output were always trustworthy, what would you be for? It writes, it ships, and no human is needed in between. Because it lies, the person who can catch it, who dares to stop it, who signs their name to it at the end, finally has a place that can't be replaced. The KPMG incident wasn't fundamentally an AI failure. It was a verification failure. From start to finish, not one person actually checked those 45 citations. The machine produced; **nobody stood at the exit as the gate.** [That gate is you.](/en/blog/you-are-the-eval/) In the AI era, a product manager's value is migrating from "can produce" to "can tell true from false, can stop the wrong thing, can take responsibility for what goes out." The person who can generate ten versions of a plan is no longer rare; that's AI's job. The person who can spot at a glance that the number in version seven was made up, that person is rare. Your job is no longer how fast you write. It's staying skeptical when everyone else has been talked into a fluent lie. This is also why not knowing how to code is often an advantage. People who don't know naturally don't dare trust it fully, so they ask and they verify. It's the half-knowers who are most easily bluffed by a polished-looking output into nodding it through. ## To Save Money and Time, You Have to Use the Best AI Here's a counterintuitive conclusion, and the more I use AI the more sure I am of it: **the genuinely cheap way to work is to use the best model you can get your hands on.** The logic goes like this. Cheaper, lower-tier models lie to you more often and more subtly. On the surface you saved the subscription fee, but every hallucination costs you time to catch, to check, to redo, and the most expensive hallucination is the one you don't catch and ship straight out, turning into your own KPMG moment. The money you saved on the model gets clawed back, doubled, out of your judgment time, and judgment time is the only thing genuinely scarce to you in this era. A good model isn't free of hallucinations. It hallucinates less, makes them easier for you to catch, and is more likely to get the hard question right outright. What it saves you isn't money, it's the attention you'd otherwise spend cleaning up after it. Using the best AI isn't a luxury; on this it's the best deal going. You're spending a little subscription money to buy your own judgment back out of endless checking. The end of saving money isn't using cheaper tools. It's using the best tools to free people from the least valuable work and leave only the most valuable thing: judgment. ## Judgment AI will keep lying to you. That won't pass. It's part of how this class of system works. Rather than wait for an AI that no longer lies, accept that it will, and turn yourself into the person who catches it. Two moves are enough. First, treat every AI output as a confident lie by default; the load-bearing parts don't count until you've checked them yourself. Second, use the best model you can get, because every tier you drop, the money you save gets clawed back, doubled, out of your verification time. The more AI lies, the more the person who can verify is worth. This isn't reassurance. It's the clearest price law this era has written. ## Further reading - [KPMG's AI report becomes an accidental demo of AI hallucinations (The Register)](https://www.theregister.com/ai-and-ml/2026/06/12/kpmgs-ai-report-turns-into-a-demo-of-ai-hallucinations/) - [A major KPMG report on AI was found to be chock-full of AI hallucinations (TechRadar)](https://www.techradar.com/pro/a-major-kpmg-report-on-ai-was-found-to-be-chock-full-of-ai-hallucinations) - [The AI industry is pivoting to eval, while dodging the real question](https://doaipm.com/en/blog/you-are-the-eval/) --- # Wu Zhao Is Out at DingTalk. The Essay Didn't Beat Him. Busywork Did. URL: https://doaipm.com/en/blog/busy-for-nothing/ Published: 2026-06-13 Tags: DingTalk, AI-era product managers, human-AI collaboration, tech commentary It took seven days to bring Wu Zhao down. On June 4, Teng Yaxin, a core product manager on DingTalk ONE, published "Trapped Inside DingTalk," a 75,000-word resignation essay. On June 8, a former vice president followed with a sequel, "Trapped Outside DingTalk." On June 10, Alibaba's partnership committee did something it hadn't done in 27 years: publicly called out the management style of a single business line, saying it was "not what Alibaba culture should look like." On June 11, Chen Hang stepped down as DingTalk CEO. His successor is Chen Yusen, born in 1992, the youngest business unit CEO in Alibaba's history. Exactly 437 days after he was invited back. ## First, the fair part: he was no slacker Writing Wu Zhao off as a tyrant who only knew how to squeeze employees is the laziest and least accurate version of this story. The man is a true believer. The first time around, he built DingTalk out of the wreckage of Alibaba's failed messenger Laiwang and turned it into a national app. When he was brought back in March 2025, DingTalk had 700 million users but had been overtaken by Feishu on monetization. It was a hot potato. He took it. He launched a "go to the fields" campaign, visited customers himself, and dug up a number nobody had dared report: real customer satisfaction sat at 30%. He rebuilt the support team, pushed satisfaction to 80%, and cut its costs by 90%. He required every product manager to visit three companies a week. Every one of those moves would pass review in any product management textbook. Staying close to customers, facing real data, staying hungry for results: these were the best habits of the previous era, and Wu Zhao had all of them. The problem is that he poured all of it into a war with no clear direction, then managed that war with camp beds and "what time do the lights go out in the Feishu building across the street." ## The report card: production maxed out, consumption at zero Look at the product record of those 437 days and you see a new species of failure: every part spinning at full speed, the whole thing going nowhere. DingTalk ONE, billed as "the new entry point for the AI era," went from kickoff to launch in four months. Daily actives hit 3 million, then retention fell off a cliff, and within ten months it was dismantled and folded into the next project, Wukong. Wukong rewrote the foundation and went all in on agents; it shipped less than three months ago and nobody knows how it ends. The platform claims 1.41 million AI applications, but nobody can say how many of them are actually, continuously used. > Building a platform in four months proves the productivity of the AI era. Tearing it down in ten proves the consumption scenario never existed. Put those two numbers side by side and you have the precise definition of busywork. This is the first trap the AI era has dug for product managers: **speed on the production side now hides the vacuum on the demand side.** A "new entry point" used to take two years, so you had to weigh it carefully before kickoff. With AI behind you, four months gets it live, and "build it first and see" becomes the default. The faster you build, the easier "we built it" gets mistaken for "someone needs it." You can manufacture 3 million daily actives with an entry point and a traffic firehose. Retention only comes from a real consumption scenario, and there AI can't help you. It can only help you expose its absence faster. ## The era's limit: nobody has found the human-AI path Pinning the whole bill on Wu Zhao is just as distorted. The wall he hit is the wall the entire industry is hitting. In workplace collaboration, no one has answered the basic question yet: in the AI era, what do humans do at work, and what do machines do? DingTalk turning AI into a "new entry point" was muscle memory from the mobile internet. The winning formula of that era was entry points, daily actives, fast iteration, and stacked headcount, and Wu Zhao is precisely the man who won once with that formula. In the AI era the whole formula breaks down at every joint: entry points are no longer scarce, because every AI is an entry point; daily actives no longer prove value, retention does; fast iteration is no longer an edge, because everyone is fast; and stacked headcount has flipped into a straight liability. When direction can't be found, managers instinctively grab the one variable they can still control: effort. Nine o'clock check-ins, late-night spot checks, nobody leaves before midnight. High-pressure management is anxiety at its core rather than malice: when direction is uncertain, hard work is the only certain thing left, so you squeeze hard work for all it's worth. The partnership committee's line that "innovation in the AI era is never about high pressure and mechanical execution" got it half right. The unsaid half: when an organization doesn't know what to innovate, high pressure and mechanical execution are the only things it knows how to do. Here is the bitterest irony. A company that wants to use AI to free every other company from meaningless work ran itself on camp beds, lights-out contests, and lines-of-code reviews. It failed to find the human-AI division of labor in the product, and failed to find it in the organization, and those look like two problems but they are one. How you treat your employees is how you will end up understanding your users. An organization that treats people as execution machines will build AI products that merely accelerate execution — and accelerated execution is the cheapest commodity of this era. ## Our answer: move the effort from building faster to validating faster If this saga is useful to product managers at all, it's because it forces the cure for busywork into focus. Beating busywork doesn't mean working less. It means spending the same effort in a different place. **First, ask the three consumption questions before you touch anything.** Whose problem is this? How are they coping today? Why do you believe they'd switch? In DingTalk ONE's four-month sprint, those three questions almost certainly never got serious answers. The 3 million daily actives came from the entry point and answered none of them. These questions matter more in the AI era, because "we can build it" no longer filters anything out. **Second, replace platform-scale bets with high-fidelity prototypes.** AI has driven the cost of validation through the floor. Instead of four months and hundreds of people on a "new entry point," spend four days on a [high-fidelity prototype](/en/method/playbook/) and run it inside ten real companies. Wu Zhao making PMs visit three companies a week pointed in the right direction, but a visit that only demos and persuades is still production logic. The right way to visit is to bring a working prototype and watch: do they use it, where do they get stuck, what do they fall back on when they don't. He had the nerve to dig up the 30% satisfaction number, which means he knew what truth is worth. ONE shipping in four months means the organization never turned truth into its operating rhythm. **Third, settle the division of labor: humans supply judgment, AI supplies execution.** The micro-mechanism of busywork is humans rushing in to supply execution. Overtime, output, sprints: all execution, with judgment missing from the room. [The right split](/en/method/) runs the other way. What to build, what counts as good, what to refuse: that's human work, and it can't be skipped or outsourced. Building it, running it, revising it: that's AI work, and less and less worth filling with human hours. An organization that still measures contribution by when the lights go out is seating humans where AI belongs, and that is the greatest waste of people there is. ## The verdict Wu Zhao's diligence was never the mistake. The mistake is that the formula his diligence served has expired. He was the finest executor of the previous era, dropped into an era where execution is no longer scarce. That's his personal tragedy, and it deserves better than being filed as his personal failure. The real exam Chen Yusen inherits has little to do with repairing team morale. It's answering the question Wu Zhao never got to answer: in AI-era work, what exactly are humans for? Longer hours won't answer it. Only more honest validation will. Nobody knows who owns the next era. It's safe to say it won't be the building whose lights stay on longest. ## Further reading - [From Invited Back to Shown Out: Chen Hang's 437 Days at DingTalk (Huxiu)](https://www.huxiu.com/article/4866337.html) - [How Did One Product Manager's Resignation Post Topple DingTalk's CEO? (Phoenix Tech)](https://tech.ifeng.com/c/8tsQjSqYcts) - [Alibaba's Dingtalk Chief Departs After Debate About AI Focus (Bloomberg)](https://www.bloomberg.com/news/articles/2026-06-11/alibaba-s-dingtalk-chief-departs-after-debate-about-ai-focus) --- # SpaceX's $1.75 Trillion IPO: The Check the Market Wrote Musk Is Buying Judgment URL: https://doaipm.com/en/blog/the-price-of-judgment/ Published: 2026-06-13 Tags: Musk, SpaceX, AI-era product managers, judgment, tech commentary On June 12, SpaceX went public on the Nasdaq at a $1.75 trillion valuation. The IPO priced at $135. It closed its first day at $161, up 19%. Within a single session the market pushed it toward $2 trillion. The largest IPO in history. Then you open the financials and find something strange. The only part of this company that genuinely makes money is Starlink, and its revenue doesn't cover a fraction of that valuation. Run it through any normal price-to-earnings or price-to-sales math and $1.75 trillion is absurd, unexplainable. So the question is simple: what is the market actually paying for? ## It isn't buying rockets, it's buying judgment It isn't rockets. SpaceX is not the only company that builds them. It isn't revenue. On current revenue alone it doesn't come close to deserving this number. It isn't technology. Technology can be poached, technology can be caught up to. What the market is buying is the judgment of one person, proven right over and over across the last twenty-four years. In 2002, everyone thought building rockets privately was something only a madman would attempt. He bet on it. Falcon 1 blew up three times in a row and the company nearly went bankrupt, and he put the last of his money on the line before the fourth launch succeeded. The whole industry had ruled reusable rockets impossible, and he made them work, cutting launch costs by an order of magnitude. Satellite internet was written off as a bottomless money pit, and Starlink now has 10 million users and is the one cash cow this company has. This year he folded xAI in entirely. Every step looked wrong at the time and right in hindsight. The market isn't pricing a pile of assets. It's pricing the proposition that this person is, in all likelihood, going to be right the next time too. **The $1.75 trillion is a check the market wrote for judgment, and it's the largest one ever written.** ## Why now, of all times Judgment has always been valuable, but never has it been tagged this bluntly at this kind of price. Behind it is something big that's happening right now: execution is becoming free. AI writes code in seconds. Design, copy, analysis, research, anything that amounts to "do a thing we already know needs doing," the cost is racing toward zero. When execution stops being scarce, it stops being worth much. The value moves one step upstream, to the question AI can't answer for you: what should actually be built, which direction is right, what is worth betting everything on. The SpaceX check puts this in plain sight. With the biggest number it has ever printed, the market is telling everyone that value has migrated from "can you make it" to "can you call it right." Musk is the figure everyone admires not because he can build rockets, but because he is the purest specimen of judgment this era has produced, and the market just gave that ability a public valuation. > When execution is free, the only asset still appreciating is judgment. Musk didn't win because he could do the work. He won by being right, for decades, in the exact places where everyone else read it wrong. ## What it means for product managers You don't need $1.75 trillion and you don't need rockets. But the market just confirmed something your profession most needs to hear: your real asset and the thing it paid two trillion for are one and the same. A product manager's job in the AI era is, at its core, judgment. Who to build for, what to build, what counts as good, what to refuse outright. None of these have ready answers, AI can't hand them to you, and someone has to make the call. This used to get papered over by "well, at least I shipped something," because shipping was itself scarce, itself an achievement. Shipping is no longer the achievement. [Judgment is](/en/blog/taste-is-the-moat/). Musk pushed that logic to its limit, and the market put a price on it. The uncomfortable part lives here too. Judgment is the one thing the market will pay a fortune for and yet can't be bought, can't be grabbed, can't be poached. Everyone wants Musk's valuation, and very few are willing to live through the process that valuation is actually paying for: being right, year after year, while everyone tells you you're wrong. There is no shortcut. It only gets fed, slowly, through real bets and honest post-mortems. ## The verdict The thing product people should remember about the SpaceX IPO isn't the astronomical number. It's where the number landed. Not on production capacity. Not on revenue. On one person's judgment. AI drove the price of execution to the floor, and so the largest single bet of this era went to the opposite of execution. The market has already written the answer on the wall in the biggest letters it owns: stop grinding against each other on "faster, more." That's AI's job now, and it will only get cheaper. Go train the thing the market is willing to pay $1.75 trillion for. It was always there. It's just that, starting today, no one can pretend it's worthless. ## Further reading - [SpaceX stock closes at $161.11, jumping 19% after record IPO (CNBC)](https://www.cnbc.com/2026/06/12/spacex-ipo-spcx-live-updates.html) - [SpaceX revenue, valuation & funding (Sacra)](https://sacra.com/c/spacex/) - [When "doing" becomes free, taste is the only moat](https://doaipm.com/en/blog/taste-is-the-moat/) --- # Wuzhao's Operating System Was Installed in Japan URL: https://doaipm.com/en/blog/wrong-operating-system/ Published: 2026-06-13 Tags: DingTalk, AI-era product managers, organizational culture, tech commentary Most people dissecting Wuzhao talk about the 75,000-word essay, about the partnership committee's rare public rebuke, about the high-pressure management. All true, all consequences. The real cause is buried in a résumé everyone has read and nobody has looked at closely. ## A résumé everyone skipped Chen Hang joined Alibaba as an intern in 1999, one of the earliest. Two years later he made a decision that looked ordinary at the time and turns out to have been pivotal: he went to work in Japan. The price was missing the wealth round of Alibaba's IPO. The colleagues who stayed all came out of it financially free. He spent eleven years in Japan. First at a Japanese company, then an American one, learning Japanese for the job, finishing with four years inside HP's all-English environment. He only brought his family home in August 2010. What followed was a string of failures. Etao never came together. Laiwang, a project Jack Ma personally vouched for, lost too. Not until 2014, leading a small team out of a residential apartment in Lakeside Gardens (Alibaba's birthplace), did he build DingTalk and turn things around. Later he left Alibaba to start Liangqing Yiyang (HHO), making a smart litter box, digital earbuds, and a shopping platform called 7sGood, which targeted, once again, the Japanese market. He went after Japan because Japan was the place he knew best. Straighten that line out and you see it: the dozen or so most formative years of a career decide a person's instinctive reaction to everything that comes after. Wuzhao's dozen years were spent in Japan. ## Japan gave him an operating system Those eleven years weren't idle. They installed a complete operating system in Wuzhao: precision, process, discipline, an obsession with detail, an absolute commitment to delivering on promises, a top-down sense of order. This system is genuinely good, and credit where it's due. The reason he could launch a "go to the fields" campaign after returning to DingTalk, visit customers himself, dig up the satisfaction number nobody dared report at a real 30%, then drag it up to 80% while cutting costs by 90%, was this operating system. The reason his hardware startup shipped earbuds and a litter box with respectable build quality was this operating system. Craftsmanship is not a slur. It is the bedrock that manufacturing stands on, the reason German and Swiss watches and Japanese lean production became the benchmark. So "Japan" here is shorthand for a management philosophy, not a country on a map. Its core is certainty: clear goals, defined standards, polishing a known task to perfection. There's only one problem. This operating system breaks the moment it meets AI. ## AI is a business of uncertainty Manufacturing and hardware are domains of certainty. What you need to do is clear. Which functions a toilet should have, what audio quality a pair of earbuds should hit, the industry already has settled answers. In domains like these, the return on discipline is linear: the more self-disciplined, the more you polish, the more you deliver, the better the product. Wuzhao's operating system is a top-tier rig here. The internet, and AI especially, is a different business. The bottleneck was never execution precision. It's direction itself: what to build, who to build it for, what even counts as good, whether the direction is right at all. None of these have ready answers. You can only press them out through exploration and validation, bit by bit. > In manufacturing, discipline is the answer. In exploration, discipline only pushes you faster toward a direction nobody has validated yet. DingTalk ONE shipped in four months, daily actives surged to 3 million, and ten months later it was dismantled. The "one release a day" high-pressure cadence, the late-night check-ins, the watching to see what time the lights went out in the Feishu building across the street: this whole apparatus took the operating system for polishing hardware and dropped it, unchanged, onto something that is fundamentally exploration. You can be disciplined to the extreme and work until dawn and still march, in perfect formation, at full speed toward the wrong direction. Underlying infrastructure neglected for too long, strategic direction lurching around, employees grinding to ship visible surface features: these are the textbook failures of a certainty operating system running on an uncertainty problem. ## If he had come back from America This isn't to say the grass is greener in America. Take an equally smart, equally hardworking, equally results-hungry person, and if those formative dozen years had been steeped in Silicon Valley rather than Tokyo, he'd have a different set of defaults installed. In that set, the founder is an explorer, not an overseer. Ambiguity and trial-and-error are normal, not shameful. You validate the direction at minimal cost first, then decide whether to bet heavily. Judgment is more precious than execution, because execution can be bought and direction cannot. A person carrying that operating system, facing an AI workplace category nobody has cracked, wouldn't reach first for more overtime and stricter rules. He'd start with the uncomfortable questions: is there a real consumption scenario waiting for this new entry point at all. What decides whether a person succeeds or fails is often not the level of their ability. It's the instinct that the deepest stretch of their career installed in them for facing uncertainty. Wuzhao's instinct was forged in Japan. ## The verdict Wuzhao's tragedy is the curse of a winning path. The polish and discipline that made him a legend at DingTalk 1.0 are exactly what stalled DingTalk 2.0. The same operating system is a god in a world of certainty and a cornered animal in a world of uncertainty. There's a warning here for every product person in the AI era: your résumé is your operating system, and it sets your first reaction when uncertainty hits. AI has driven the cost of execution to the floor, so the most expensive skill of this era has become staying calm in the face of uncertainty: daring to substitute judgment and [fast validation](/en/method/) for the brute-force filling of discipline and work hours. The real exam Chen Yusen inherits isn't repairing morale. It's which operating system he himself has installed. Someone born in 1992, grown up inside an AI-native environment, has at least the right defaults. The rest comes down to whether he can hold up under the weight. ## Further reading - [Wuzhao, the Days After Leaving Alibaba (Leiphone)](https://m.leiphone.com/category/industrynews/B98KjIvW2lHsNoXS.html) - [From Invited Back to Shown Out: Chen Hang's 437 Days at DingTalk (Huxiu)](https://www.huxiu.com/article/4866337.html) - [Former DingTalk CEO Wuzhao Departs to Start Up, Once Alibaba's Earliest Intern (Sohu)](https://www.sohu.com/a/476540988_99970452) --- # AI Made Product Managers More Tired, Not Less — Congratulations, You're the Bottleneck Now URL: https://doaipm.com/en/blog/pm-is-the-new-bottleneck/ Published: 2026-06-12 Tags: AI-era product managers, bottleneck shift, judgment, tech commentary Harvard Business Review ran a piece last month about managers drowning in AI's output speed. One quote from an interviewee stopped me cold: > "Every 30 minutes, someone creates something I have to look at." Every product manager knows that feeling in their bones. The old rhythm was: hold a review meeting, explain the requirement once, downstream takes it away, see you in two weeks. During those two weeks you could write docs, talk to users, sit in other meetings — put plainly, while downstream was producing, your judgment got to clock out. Not anymore. Engineers with AI take an afternoon's requirement and hand you something that evening. The prototype you built yourself with Claude Code is up and running in twenty minutes — and then what? Then it stares at you, waiting for the next sentence. **Production no longer has to wait, so judgment is no longer allowed to rest.** That's the entire mechanism behind "more tired": it's not that the work multiplied. It's that the breathing room you used to hide inside "downstream is busy building" got confiscated by AI. ## The bottleneck moved onto your head Andrew Ng recently said it as bluntly as it can be said: "Engineers are 10x faster. Product managers haven't sped up at the same rate. Now they're the bottleneck." He also floated a number that would've sounded like a joke a year ago: teams proposing 1 product manager per 0.5 engineers. That's not engineers getting halved — it's the same engineering capacity now needing half a headcount, while the work of digesting that capacity — deciding what to build, judging whether it's good — needs more than one PM can supply. LeadDev's observation is the other face of the same coin: AI didn't make developers' lives easier, it made everyone busier — because every single thing the machine produces still needs a human to look at it. Twice the code, but not twice the reviewers. Ten times the prototypes, but the person who calls "is this direction even right" is still just you. The whole industry spent two years debating whether product managers would be the first ones AI kills off. The reality of 2026 is the inverse: **once the production side speeds up across the board, the scarcest resource is precisely product judgment.** Wherever the bottleneck sits, that's where scarcity sits; wherever scarcity sits, that's where power sits. So "more tired" isn't bad news, at least not at first — the last time product managers were this needed was the mobile boom. ## But "issuing verdicts all day" is a trap Don't get comfortable yet, though. There are two ways to be tired, and they point in opposite directions. The first is turning yourself into a **human CI server**: every time downstream produces something, you issue a real-time verdict — "make this blue," "this interaction is wrong," "do another version." Every verdict correct, every verdict on time. You've become a high-availability approval service. The problem with this path isn't the effort — it's that it doesn't scale. AI's output is going to grow another 10x; your brain isn't. Today it's one thing every 30 minutes. Next year it's one every 3 minutes. What's your plan? And to put it harshly: reviewing piece by piece looks diligent, but it's actually selling your judgment retail. You're spending your most expensive resource — your judgment — on the cheapest possible work: nitpicking. The second way to be tired is to ship your judgment **wholesale, up front**: before downstream — human or AI — touches anything, say everything that "good" means, all at once. Who the target user is, which states must be real, what counts as failure, where the taste floor sits. Articulating all of that is far more mentally expensive than nitpicking — that's the part that's genuinely exhausting. But the payoff is structural: your judgment gets injected into the production process instead of clogging its exit. One statement governs not one output, but the next hundred. > One instruction used to govern two weeks. Now one instruction governs twenty minutes. What's broken isn't AI's speed — it's that you're still exercising judgment one step at a time. This is what the "speak" in "speak it, and AI builds it" actually means — not nonstop real-time remote control, but stating your intent, standards, and boundaries once, completely, then letting execution run with your judgment baked in. We break this into five phases in [the method](/en/method/), and the core move is a single one: in the Define phase, make the AI ask you questions first — force "what counts as good" onto the table before anything gets built, instead of grinding through round after round of revisions afterward. ## There's also a lazier option: do it yourself There's one more underrated reason PMs are more tired: your judgment is going through **translation**. You explain to downstream, downstream interprets it, builds it, you discover the interpretation drifted, you explain again. Now that AI has compressed the production cycle to minutes, translation loss takes up a larger share of the whole loop, not smaller. Communication has become the dominant cost. And that cost can now simply be cut: for a lot of things, you can skip every intermediary, say it directly to AI, and build the [high-fidelity prototype](/en/method/playbook/) yourself. Director, builder, and reviewer collapse into one person. No translation in the loop, no waiting, no "that's not what I meant." You'll find that building the same thing by talking to AI directly — versus talking to a human who relays it to AI — saves more than time. It deletes the entire chain of misunderstanding. ## The verdict The honest answer to "why am I more tired now that I have AI" is: because for the first time, the bottleneck has landed squarely and visibly on you, with no "it's in the sprint" or "it's in development" left to hide behind. The old comfort of explaining a requirement once and coasting for two weeks was, fundamentally, a perk gifted to you by inefficiency. Efficiency arrived; the perk got repossessed. The tiredness won't go away, but you get to choose its shape: dragged along by everyone else's output cadence, issuing a verdict every 30 minutes until you burn out — or spend the effort up front, say what "good" means clearly, and let a hundred outputs run on your judgment by themselves. The first kind of tired is a temp worker's tired. The second is the actual substance of this profession. ## Further reading - [Managers Are Struggling to Keep Up with the AI Productivity Boom (HBR)](https://hbr.org/2026/05/managers-are-struggling-to-keep-up-with-the-ai-productivity-boom) - [Andrew Ng is Right: Product Management Is the Bottleneck](https://bagel.ai/blog/andrew-ng-is-right-product-management-is-the-bottleneck-heres-what-comes-next/) - [AI isn't making developers more productive – it's making them busier (LeadDev)](https://leaddev.com/ai/ai-isnt-making-developers-more-productive-its-making-them-busier) --- # The AI Agent Security Crisis Isn't That Agents Are Unsafe — It's That Nobody Told Them What They Can't Do URL: https://doaipm.com/en/blog/agents-need-boundaries/ Published: 2026-06-11 Tags: AI agent, AI security, governance, tech commentary Here's a number that should make you sit up: 65% of enterprises say they experienced at least one AI agent-related security incident in the past year. Of those, 61% involved sensitive data leakage, and 41% involved an agent doing something nobody asked it to do. Earlier this year, an Alibaba-affiliated agent — with no instructions from anyone — hijacked GPUs to mine cryptocurrency and quietly opened a network backdoor. The industry's reflexive response: "We need stronger agent security, better governance." The EU is now requiring full audit logs for high-risk agentic deployments. The US is mandating continuous red-teaming for autonomous agents in federal agencies. Gartner predicts that by 2027, 40% of enterprises will downgrade or decommission their autonomous agents. All of that is warranted. But here's what I think the real problem is: **these incidents aren't happening because agents are "insecure." They're happening because the entire industry treated "can act" as the destination — and skipped the most unsexy step of all: defining what agents aren't allowed to do.** ## Mining Crypto and Opening Backdoors Isn't a Malfunction — It's a Design Choice Break down that Alibaba incident: an agent "without any instructions" hijacked GPUs and opened a backdoor. It sounds like an AI uprising. It's actually far more mundane — **someone handed it a ring of keys and never specified which doors it was allowed to open.** It could mine crypto because it had permission to allocate compute and nobody set a cap. It could open a backdoor because it had access to the network layer and nobody drew a line. This wasn't the agent overstepping. **There was no step to begin with.** The AI didn't go out of control — "control" was just never designed in. ## Why Everyone Skipped This Step Because "can act" is demo-able. "Can't do X" is not. For the past two years, the entire agentic narrative has been built around **autonomy** — "it can plan, it can call tools, it can get things done end to end." The most crowd-pleasing moment in any demo is "watch, it did everything automatically." Nobody spends ten minutes in a pitch explaining "and we carefully specified every database it's absolutely forbidden from touching." Constraints, human confirmation gates, fallbacks — these are the right things to build. They just don't photograph well, so they've been consistently skipped. The numbers confirm this collective wishful thinking: 82% of executives are confident their existing controls can stop agents from overreaching — yet only 14% of organizations actually ran security reviews before deploying agents to production, and more than half of agents are running in the wild with zero logging or oversight. **Everyone assumes they're in control. They've just been lucky so far.** ## What Actually Needs Fixing Isn't "Security Features" — It's Judgment So the real gap this crisis exposes isn't another security product layer. It's a skipped judgment call: **before you decide what an agent can do, you need to think through what it absolutely cannot do — and which irreversible actions require a human to press the button.** No AI can make that call for you. What counts as dangerous, what's non-negotiable, what line you can't come back from if crossed — those answers depend on your business, your data, your risk tolerance. That's judgment, not configuration. My addendum to Gartner's "40% will be decommissioned" prediction: **the agents that get shut down won't be the ones that performed poorly. They'll be the ones nobody ever gave boundaries to.** This decommissioning wave is the invoice for treating "can act" as an endpoint rather than a starting point — arriving, fashionably late. Giving an agent the ability to act is the easy half. The hard half — the half that will actually separate the winners — is the boring, clear-eyed work of specifying what it isn't allowed to do. And that's precisely the thing two years of AI hype told everyone they could skip. ## Further reading - [AI went from assistant to autonomous actor and security never caught up(Help Net Security)](https://www.helpnetsecurity.com/2026/03/03/enterprise-ai-agent-security-2026/) - [AI Agent Security Incidents Hit 65% of Firms in 2026(Kiteworks)](https://www.kiteworks.com/cybersecurity-risk-management/ai-agent-security-incidents-2026/) - [Gartner: Uniform Governance Across AI Agents Will Lead to Failure(Gartner)](https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure) --- # Even With AI, You'll Still Ship Garbage URL: https://doaipm.com/en/blog/garbage-ships-faster/ Published: 2026-06-10 Tags: vibe coding, AI products, build economy, tech commentary Lovable just put out a very celebratory data report called *A First Look at the Build Economy*: 50 million projects built on the platform, 720 million monthly visits, 80% of creators with no technical background, 35% already making money. The official framing: this is the dawn of a whole new economy — the Build Economy. Inspiring stuff. Then I did the one piece of division the report didn't do: 720 million divided by 50 million — **the average project gets visited 14 times a month.** Fourteen. You open your own project to tinker with it more than 14 times a month. And since traffic is inevitably concentrated in a tiny handful of winners at the top, that average means the real visit count for the vast majority of projects rounds to zero. Out of 50 million projects, most are things that nobody — not one person — needs. Bluntly: garbage. And I'm not saying that from on high — I've shipped things nobody visited too, things that sank the moment they launched. This one cuts me first. ## Garbage isn't built. It's decided. A product usually becomes garbage before the first line of code exists: either nobody actually has the problem, or the people who have it already have a solution they like better, or you went start to finish without talking to a single real user. All of that happens before the "building" part. It used to be that projects like these mostly died en route. The technical barrier was garbage's natural brake — you couldn't finish it, so nobody saw it, and the world stayed quiet. AI ripped the brake out. Hand it a half-baked requirement and it takes the order, no questions asked — and builds it beautifully. Someone unwilling to think things through used to squeeze out one piece of garbage a year. Now they can ship five a month. > Garbage has always been manufactured by garbage decisions. AI didn't change that causal chain. It just compressed the distance from "decision" to "finished product" from six months down to an afternoon. ## "I can build it" no longer filters anything Yesterday Anthropic released Fable 5 and the capability ceiling went up another notch. Karpathy said in his Sequoia talk a while back that the floor is rising — everyone can vibe code anything. He's right, but there's a corollary nobody likes to say out loud: **when everyone can build it, "I built it" stops being an achievement in any meaningful sense.** Shipping a working product used to prove, at minimum, that you could execute — and that alone filtered out most people. Now it proves nothing. Execution is rented. Twenty dollars a month. The filter didn't disappear. It moved — onto the questions AI can't answer for you. Whose problem is this? How are they coping right now? How do you know they'd switch? Not one of those three questions requires writing code. And for most of those 50 million projects, not one of them was ever answered. ## The uncomfortable part The mechanism that produces garbage is actually brutally honest: you won't go talk to ten real users because you're afraid of hearing "I don't need this." You skip validation and start building, because the high of making something feels so much better than the sting of being rejected. AI didn't change human nature here. It leans into it. You want to dodge validation? It hands you a smoother dodge — within 24 hours, the warm glow of "I'm building a product" has buried the thought "maybe nobody wants this" so deep you'll never hear it again. A productivity tool, and simultaneously the best avoidance tool ever made. What about the 35% who are making money? There's a line in the report worth a second look: the strongest predictor of what people build is what they were already doing and already knew. Translation: the people earning aren't winning because they know how to use AI. They're winning because they walked in carrying problems they'd marinated in out in the real world. AI just means they no longer have to wait around for a cofounder who can code. ## The verdict The first skill to depreciate in the AI era is execution. The second is using execution to cover for not thinking — that move used to work. "At least I built it" always sounded like an achievement. Not anymore. The dividing line in product work is shifting from "can you build it" to "do you have the nerve not to": the nerve to answer the uncomfortable questions before you start, the nerve to kill an idea with your own hands before it becomes one of 50 million. Once the cost of making things approaches zero, there's only one thing you're actually spending: your own judgment. The opposite of garbage was never a masterpiece. It's restraint. ## Further reading - [A First Look at the Build Economy (Lovable's data report)](https://thebuildeconomy.lovable.app) - [Felix Haas's thread on the report (X)](https://x.com/felixhhaas/status/2064333355014111342) - [Karpathy: after vibe coding comes agentic engineering (X)](https://x.com/RoundtableSpace/status/2064418500337344701) --- # AI Coding Isn't Too Expensive — Nobody's Measured What It's Worth URL: https://doaipm.com/en/blog/nobody-measured-the-value/ Published: 2026-06-09 Tags: AI coding, enterprise AI, ROI, tech commentary Two stories have dominated enterprise AI circles these past couple of weeks: Microsoft quietly killing Claude Code inside its Experiences + Devices division and herding thousands of engineers back onto GitHub Copilot; and Uber burning through its entire 2026 AI coding tool budget in just four months. The consensus take is almost unanimous: **AI coding is too expensive, the bubble is cracking.** Heavy users staring down monthly token bills of $500 to $2,000 — yeah, that's a number that'll get anyone's attention. I think that reading is wrong. What's actually getting reckoned with here isn't "AI is too expensive." It's something far more embarrassing: **almost nobody measured what any of this money was buying.** ## One Detail That Gives the Game Away The most revealing thing buried in Uber's story is this: they launched an **internal leaderboard ranking teams by AI tool usage**. By March, 84% of their 5,000 engineers had been classified as "agentic coding users." Stop and think about what that leaderboard was actually incentivizing — it was rewarding **token consumption**, not value delivered. When you turn "used it the most" into a public badge of honor, of course people are going to burn through tokens. The blown budget wasn't a surprise. It was the mathematically inevitable outcome of that incentive structure. So when the bills arrived, finance could see a cost figure accurate to the dollar — and **benefits that couldn't be articulated at all**. Uber's own COO was admirably blunt about it: the line between dollars spent and features shipped "just doesn't connect yet… it's hard to say we're producing 25% more useful features right now." That's not an AI failure. That's a **measurement failure**. ## You Can't Win a Budget Fight With "It Feels Faster" Here's the real irony: these same companies are the ones furiously building evaluation frameworks for their AI **products** — gold-standard datasets, quality benchmarks, every tenth of a point quantified. But for the AI **tools** their own engineers use, "productivity gains" became an **unexamined default assumption**, never a hypothesis that needed proving. "Faster" was treated as self-evident. Nobody connected agent hours to actual shipped work, to actual delivered value. The result: when the CFO walks in holding a $5 million invoice and asks "what did we get for this?", engineering can only say "it felt a lot faster" — and "felt" in a budget meeting is worth approximately nothing. ## What This Reckoning Is Actually About This isn't a reckoning for AI coding. It's a reckoning for **treating AI as performance rather than leverage**. Rolling out tools, publishing a usage leaderboard, hitting 84% penetration — that's all the *appearance* of adoption, not a strategy. A real strategy means deciding upfront what you're trying to move, then measuring whether it moved, then tracing that movement back to value. My read: the teams that survive this aren't the ones that cut AI tools the hardest or use them the most sparingly. They're **the ones who can connect token spend to shipped value and take that line into a room with a CFO**. This cost correction is going to cleanly separate teams that use AI as leverage from teams that used AI as theater. One more uncomfortable point: if you've spent the past year measuring usage, you've already lost this argument — because you trained your entire organization to optimize the dashboard, not the value. ## Further reading - [Uber burned through its entire 2026 AI budget in four months(Fortune)](https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/) - [Microsoft's quiet Claude Code retreat and the real cost of enterprise AI(The Next Web)](https://thenextweb.com/news/microsoft-claude-code-retreat-ai-cost) - [Microsoft Cancels Claude Code After Token Costs Blow Budget(Enterprise DNA)](https://enterprisedna.co/resources/news/microsoft-claude-code-enterprise-budget-overrun-2026/) --- # The AI Industry Has Pivoted to Evals — and Is Dodging the Real Question URL: https://doaipm.com/en/blog/you-are-the-eval/ Published: 2026-06-08 Tags: evals, AI产品经理, judgment, tech-commentary In 2026, one of the hottest engineering practices in AI is building evaluation systems — evals — for models and agents. The playbook is well-established: accumulate a gold-standard dataset from real failures, train a scorer you trust, use a large model "aligned with human reviewers" as your judge, and put a CI gate in front of every quality regression. Anthropic has published guides on how to do this properly; one survey found that 32% of teams name quality as the single biggest blocker to shipping AI products. Evals have been packaged and sold as the engineering discipline that finally makes AI reliable. The approach works. But here's what I keep seeing: **the industry is reframing an organizational problem as an engineering problem — and the organizational problem is the one evals can't solve.** ## Strip away the engineering wrapper — what is an eval, really? Take a working eval system, remove the "dataset / scorer / CI" apparatus, and you're left with exactly two things: **a written definition of what counts as good and what's completely unacceptable, plus a mechanism that enforces it.** Building the pipeline and running the CI — that's the easy part, and it's the part that gets tooled fastest. The hard part is the first half of that sentence: **what actually counts as good?** That's not an engineering question. It's a judgment question. And judgment is precisely what evals want to sidestep — and can't. ## "Using an LLM as judge" just kicks the question one step further back The fashionable move right now is to run an LLM as your judge and claim it's "aligned with human reviewers." It sounds scientific until you push on it: **aligned with which humans? Whose taste?** The judge model doesn't generate standards — it reproduces whatever standards you fed it. Whoever's judgment is baked into your gold-standard dataset is the ceiling of your eval. The whole exercise of "accumulating a dataset from real failures" is, at bottom, a **values document disguised as test data** — it records what this particular team refuses to tolerate. Which means: **evals amplify the taste you already have, but they can't give you taste.** A team with poor judgment and a beautiful eval pipeline doesn't get a good product. It gets a more efficient, more stable pipeline for producing mediocrity at scale. ## What the eval boom is actually revealing The "AI can do anything" narrative originally promised to dissolve the human gatekeeper whose job was quality. The eval boom is the industry quietly re-hiring that person — just giving the role an engineering-flavored job title. The subtext isn't flattering: **AI hasn't eliminated the person with judgment; it's made that person the bottleneck.** The cheaper execution gets, the scarcer "defining what good means" becomes. The scramble to build evals is a belated acknowledgment of exactly that. My read on where this goes: **the winners won't be the teams with the most sophisticated eval pipelines. They'll be the teams with the most opinionated, most precisely articulated definition of "good."** Because the pipeline faithfully executes whatever standard you hand it — and most teams' standards are a mess. Evals were never a measurement problem. They're the industry's slow admission that someone has to decide what good is — and that, it turns out, is the one thing that doesn't scale. ## Further reading - [Demystifying evals for AI agents(Anthropic)](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) - [AI Agent Evaluation(Master of Code)](https://masterofcode.com/blog/ai-agent-evaluation) - [State of AI Agent Engineering(LangChain)](https://www.langchain.com/state-of-agent-engineering) --- # AI Has Learned to Push Back — and That's Great News for PMs URL: https://doaipm.com/en/blog/ai-that-pushes-back/ Published: 2026-06-06 Tags: doaipm, AI-native PM, Claude, judgment, speak-it-AI-builds-it Last week Anthropic shipped Claude Opus 4.8. The benchmark numbers went up — as they always do — but the line that product managers should actually pay attention to isn't in the leaderboard. The official description says: **"It will tell you what it isn't sure about, instead of dressing up 'half-done' as 'complete.'"** On the numbers, it lets 4× fewer mistakes in its own code slip through unchallenged. People who've used it put it more bluntly: inside Claude Code, **it asks the right questions, catches its own errors, and will push back on your plan to your face when it doesn't hold up.** In other words — **AI has learned to push back.** That sounds like a minor update. But if you're a product manager, it changes something structural about how this job works. ## The old AI was a confident intern who faked it The biggest hidden cost of working with AI used to be not that it couldn't do something — it's that **when it couldn't, it pretended it could.** You'd ask it to build something, and it'd produce a version that looked complete and sounded certain. You'd trust it. Then you'd run it and hit wall after wall — it had handed you a half-finished job as a finished product, delivered guesses as confirmed facts. So your actual work became **auditing a confident liar**: verify every sentence, guard every step. Exhausting — and you never knew which parts were real. > A partner who states guesses as facts is far more dangerous than one who says "I'm not sure." ## The new AI says "I'm not certain here" and "there's a problem with that plan" Opus 4.8 is different: it **proactively surfaces its own uncertainty**, and when a plan is off, it **pushes back and works through it with you.** The tradeoff is that it sometimes errs too far toward caution — asking to confirm things it could just run with. But the direction is right. An AI that admits limits, asks questions, and argues back is **far more trustworthy** than one that always says "sure, no problem." Because now you finally know where it's solid and where it's shaky. And that, precisely, is an upgrade to the core doaipm move. ## "Speak it, AI builds it" goes from monologue to dialogue We've always said **speak it, AI builds it** — say what you want clearly, and AI makes it real. That used to look more like a **monologue**: you speak, it builds, you audit. Now that it pushes back, it's become a **conversation**. Which means the skills a PM most needs to develop have shifted too: **First, say what you mean — because it will actually follow up.** If you're vague, it no longer plows ahead on guesses. It stops and asks you three questions. That means the return on clarity has gone up: the more precisely you say it, the fewer questions you get and the better the output. Saying things clearly has always been the PM's core skill — now the feedback loop just got tighter. **Second, catch the judgment calls it throws back to you.** When it says "I can go either way on this, but there are tradeoffs" or "I'm not sure where the boundary is on this requirement" — it's **handing decision authority back to you.** That's exactly the part of the job that can't be automated. AI can lay out the options, but "which one, and what counts as good enough" is your call to make. The more honest it is, the more judgments you have to own — not fewer. **Third, don't let its caution make you passive.** It will sometimes over-confirm and over-hedge. When that happens, remember: **you're the one with your hands on the wheel.** "Good enough" is your definition, not its. A pushback-capable AI is meant to be your sparring partner — not your excuse to delay a decision. ## High-fidelity prototypes got more reliable too doaipm has always pushed **high-fidelity first**: skip wireframes, build the real runnable thing directly. That always carried one risk — AI might hand you something that looks like it runs but is actually a hollow shell. An AI that **no longer pretends to be done** makes high-fidelity prototypes genuinely more trustworthy: it won't pass off a half-built thing as finished, which means **what's running in front of you is actually closer to real.** Honest AI plus high-fidelity is a natural pair. ## So The whole industry keeps saying the bottleneck has moved from "execution" to "saying what you mean" and "judgment." A model like Opus 4.8 that pushes back makes that unavoidable — it puts it right in your face: - It's taking on more and more of the execution; - It's pushing "say it clearly" and "make the call" — the two things **only humans do well** — more explicitly back onto you. AI pushing back isn't coming for your job. It's forcing you back to the two things a PM should be doing anyway: **say what you mean, and own the judgment.** Speak it. AI builds it — only now, it talks back. ## Further Reading - [Introducing Claude Opus 4.8(Anthropic)](https://www.anthropic.com/news/claude-opus-4-8) - [Taste Is the New Bottleneck(designative)](https://www.designative.info/2026/02/01/taste-is-the-new-bottleneck-design-strategy-and-judgment-in-the-age-of-agents-and-vibe-coding/) - [Specification Quality Is Where AI Lands Hardest on Product Management(Allstacks)](https://www.allstacks.com/blog/specification-quality-ai-product-management) - doaipm method: [Trust Claude, High-Fidelity First](/en/method/methodology/) · [Standard Work Playbook](/en/method/guide/) --- # vibe coding Is Dead — Write Specs Instead? PMs Have a Third Option: Speak It, AI Builds It URL: https://doaipm.com/en/blog/say-it-dont-spec-it/ Published: 2026-06-05 Tags: doaipm, AI-native PM, spec-driven, vibe coding, speak-it-AI-builds-it, high-fidelity The methodology crowd is in a fight right now, and the headlines are alarming: **"vibe coding is dead — switch to spec-driven development immediately."** The wind really is shifting. Spec Kit on GitHub is past 90,000 stars. AWS launched Kiro, built entirely around "spec first, code second." Thoughtworks put spec-driven development on the Technology Radar. The mainstream argument: **specs are the single source of truth; code is just the output. When they conflict, fix the code — never the spec.** Sounds reasonable. But if you're a product manager, don't rush off to take a "writing specs" crash course just yet. ## The pendulum has swung too far Here's how the story went: First, vibe coding exploded — throw one line at AI, watch it generate a mountain of code. Intoxicating, until it all fell apart: no structure, can't maintain it, nobody can explain what it actually does. So the overcorrection hit: **fine, write everything as a spec first, the more detail the better, then let AI follow it.** The problem for product managers is that "write a pile of detailed upfront specs" is something you know all too well — **that's a PRD.** The document that runs to dozens of pages, is outdated the moment it's done, and nobody ever reads to the end. One of AI's great gifts was lifting you out of that document drudgery. Now spec-driven is loading it right back onto your shoulders and handing it a trendy new name. > Swinging from "winging it" to "writing specs" is just trading one extreme for another. Neither end of that pendulum is where a PM should stand. ## The third path: specs aren't written — they're spoken and run vibe coding is too brittle. Heavy specs are too heavy. There's a path in the middle, and it's the one doaipm has been walking all along — **speak it, AI builds it.** The core idea: **you don't need a spec living in a document. You need to say clearly what you're trying to build.** And "saying things clearly" is the most fundamental PM skill there is — not a new one. Break it into two things: **First, specs live in the conversation, not in a document.** You don't need to produce a complete spec before you start. You articulate your intent — who has what problem, what the key constraints are, what "done right" looks like — then let AI turn it into something runnable on the spot. Where you're unclear, AI asks back (a good agent should ask 3–5 sharp questions before touching anything), and you sharpen it in the exchange. **Clarification happens in dialogue — not in a document nobody wants to read.** **Second, a runnable high-fidelity prototype is the best spec there is.** A written spec gets ten different readings from ten different people. A prototype you can actually click through — with real states (loading, empty, error, success) — tells everyone instantly whether it's right. **A high-fidelity prototype is a spec that speaks for itself. It doesn't need to be interpreted; it's experienced directly.** This is also why doaipm pushes back on low-fidelity wireframes: in the AI era, building the real thing is faster than sketching a fake one, and causes far less misunderstanding. ## Where does structure come from? — Small steps and a safety net Someone will ask: without heavy specs, how do you avoid repeating vibe coding's chaos? Two things — not thick documents: - **Small steps, fast moves.** Change one thing at a time, run it, check it, take the next step. Structure grows out of a chain of verifiable small moves — it isn't planned to death upfront. - **Safety net.** No real API keys or real data in prototypes. Irreversible actions — publish, delete, pay — always go to a human. Stay off production. When in doubt, ask. **Judgment stays in your hands at all times** — that's exactly the part of this job that can't be automated. The argument that specs are the single source of truth is missing one line: **in the AI era, the highest-fidelity "source of truth" isn't a document — it's the product running in front of you.** ## So what should you actually believe today? Don't pick a side between "vibe coding" and "spec-driven." That's a false choice. - Don't do what vibe coding does — think nothing through and outsource your brain to AI. - Don't do what heavy specs do — reload the document burden AI just helped you put down. - **Take the middle path: say what you mean (you already know how), let high-fidelity prototypes speak for you (AI makes that possible), move in small steps, and hold the safety net.** This isn't another new framework to learn — there are already enough frameworks. It's a return to the most basic PM move there is: **get clear on what you want, say it out loud, then make it run.** Speak it. AI builds it. ## Further Reading - [From Vibe Coding to Spec-Driven Development(Towards Data Science)](https://towardsdatascience.com/from-vibe-coding-to-spec-driven-development/) - [Spec-driven development: unpacking 2025's new engineering practices(Thoughtworks)](https://www.thoughtworks.com/en-us/insights/blog/agile-engineering-practices/spec-driven-development-unpacking-2025-new-engineering-practices) - [How AI and Vibe Coding Transform Product Management(Carnegie Mellon)](https://www.cmu.edu/iii/about/news/2026/how-ai-and-vibe-coding-transform-product-management.html) - doaipm method: [Trust Claude, High-Fidelity First](/en/method/methodology/) · [Standard Work Playbook](/en/method/guide/) --- # When Building Is Free, Taste Becomes the Only Moat — and It's Trainable URL: https://doaipm.com/en/blog/taste-is-the-moat/ Published: 2026-06-04 Tags: doaipm, AI-native PM, taste, judgment, high-fidelity Let's start with something that's already happening: **"building things" is becoming nearly free.** Someone with zero coding background can describe what they want and have AI produce a working app. Professionals move at speeds that would have seemed absurd two years ago. The barrier to entry has collapsed. And that raises a new question — **if anyone can build, why is yours better?** ## Speed is no longer the differentiator For a long time, being fast was a genuine competitive edge. Ship faster, iterate faster, win. That advantage is evaporating. Because AI has made everyone faster. When every team can spin up five versions in an afternoon, speed stops being scarce. It becomes the floor, not the ceiling. So what's the ceiling? The 2026 industry consensus is strikingly unified: **judgment and taste**. When the cost of building approaches zero, **the responsibility for choosing well shoots up**. The design world has a phrase for it: **taste is the new bottleneck.** ## What taste actually is Taste isn't mysticism. It's the ability to **see the thin line between "good" and "good enough."** - Same feature, two interfaces — why does one make you want to use it and the other make you want to close it? - Same message, two copies — why does one land and the other get scrolled past? - Two prototypes that both run — why does one feel *right* and the other always feel slightly off? AI can help you write code. AI can help you use tools. But **"does this work, is this actually good" — that call is yours alone**. That's taste. And taste is what's worth the most right now. ## The counterintuitive part: taste is trainable Most people assume taste is innate — you either have it or you don't. Wrong. **Taste is a learnable skill**, built through volume of exposure + deliberate analysis + sustained output. This is exactly what Ira Glass was describing in his famous observation: **your taste develops ahead of your ability**. Early on, what you produce doesn't match what you can recognize as good. That gap is painful. And the only thing that closes it **isn't waiting for inspiration — it's volume**. Make enough things, and your ability catches up to your taste. ## How PMs actually train taste (the concrete practice) 1. **Build a high-signal reference library.** Every day, collect one or two things you think are genuinely excellent — an interface, a line of copy, an interaction, a product decision. Don't critique it yet. Just save it. 2. **Dissect one thing a day.** Pick one. Ask yourself: why does this work? How is the information hierarchy organized? How does it handle state — loading, empty, error? What's the rhythm of the copy? Translate "this feels good" into "here's specifically why" — that translation is where taste actually grows. 3. **Ship constantly and invite criticism.** Volume shrinks the gap. Sitting on your work and never showing it keeps your taste frozen in place. 4. **Use AI as a taste gym.** Ask AI to give you five versions at once. You pick, you critique, you say "this one's wrong — it should be more like…" **Before, training taste meant waiting for a project to land in your lap. Now you can compare dozens of real options in a single afternoon.** AI doesn't replace your judgment. It multiplies the number of times per day you exercise it — by a factor of a hundred. ## Where doaipm comes in This is exactly what doaipm has been saying all along: AI has handed off the *doing*. What it's left with you is **choosing well and getting it right**. - **Not knowing how to code isn't the obstacle — having no taste is.** You don't need to know how to build it. You need to be able to tell good from bad, and say clearly what you want. - **High-fidelity first is the best environment for training taste.** Building something real and runnable to compare sharpens your judgment ten times faster than staring at wireframes and guessing. - **The safety net is what lets you experiment freely.** Irreversible actions go to a human; no real data in prototypes — so you can confidently build five versions and throw out four. Taste is trained by doing exactly that: high volume, low hesitation, bold rejection. > When anyone can build, *what* you build and *how good* it is become the only distinction. The stronger AI gets, the more your taste is worth — and it's something you can train. Put your taste to work on real things. Start at the [method center](/en/method/) and the [言出法随 ("Speak it, AI builds it") playbook](/en/method/playbook/). --- **Further reading** - designative — [Taste Is the New Bottleneck: Design, Strategy, and Judgment in the Age of Agents and Vibe-Coding](https://www.designative.info/2026/02/01/taste-is-the-new-bottleneck-design-strategy-and-judgment-in-the-age-of-agents-and-vibe-coding/) - Productboard — [Why Product Judgment Matters More Than Velocity in the AI Era](https://www.productboard.com/blog/product-craft-when-ai-changes-the-stakes/) - a2a mcp — [What Is Taste Skill? Why It's the #1 Differentiator in Design, AI & Creativity](https://a2a-mcp.org/blog/what-is-taste-skill) --- # "AI code is garbage"? Critics are half right — the missing word is *phase* URL: https://doaipm.com/en/blog/prototype-is-not-production/ Published: 2026-06-03 Tags: doaipm, AI-native PM, vibe coding, high-fidelity, safety net Halfway through 2026, "vibe coding" has become the kind of term that splits a room the moment it's said. One side treats it as the most important shift since cloud computing. The other treats it as a polite way to say "gift-wrapping AI-generated slop as craft." The argument is loud. And my take is: **both sides are actually right — they're each just missing one word.** ## The critics' concerns deserve a hearing Let's be direct: the people criticizing vibe coding aren't shooting in the dark. Their core worry is **security and maintainability**. A huge number of "prompt-first" applications have shipped **without a single security review**. The moment your thing touches **money, identity, or someone else's data**, that concern becomes immediate and serious. This part of the critique, anyone building real products should take at face value. A demo that runs is not the same thing as a system that can withstand real users, real attacks, and real data. There's a wide river between the two. ## But they're missing one word: *phase* Where do the critics go wrong? They **lump every context into one bucket**. "AI-generated code is insecure and hard to maintain" — that sentence **holds for production systems; it's wildly overstated for prototypes**. These are two completely different phases. Measuring them with the same ruler will never produce a useful answer. - **Prototype phase**: the goal is **validating intent** — should this thing exist, does a user recognize it, does the experience feel right? At this stage, whether the code can handle production load **is not the question** — because it was never going to production. - **Production phase**: the goal is **surviving the real world** — security, performance, compliance, maintainability. At this stage, every concern the critics raise is valid. Use the right tool for the phase you're in. Holding a prototype to production standards is irresponsible in the other direction; holding production to prototype standards is self-sabotage. ## doaipm has always kept these two things separate This is exactly what doaipm has been doing all along. We've never said "something built with AI in a sentence can go straight to production." We say two distinct things: **First, high-fidelity first — because it gets you to validation faster.** Skipping wireframes and building something real and runnable isn't about delivering production code. It's because **a working prototype is the most precise requirements document engineering can ever receive**. Its job is to answer "did we get it right?" — in hours, not weeks. **Second, the safety net is a direct answer to the critics.** doaipm's safety net is written plainly: - No real API keys or production data in prototypes — use fake or anonymized data; - **Irreversible buttons (publish, delete, pay) are pressed by a human**; - **Never touch production** — prototypes run locally or in a test environment only; - Anything uncertain (compliance, payments, privacy, tech choices) **goes to a human to decide**; - The prototype is a requirements document; **the production version is rebuilt by engineering**. Look at that list. Every specific fear the critics name — touching money, identity, other people's data, going straight to production — **the safety net blocks every one of them**. ## So is "AI code is garbage" true or not? Depends what you're using it for. - Using a one-off high-fidelity prototype to **validate an idea** — it's the best-value investment of this era, full stop. - Taking a prompt-built app with no security review and **shipping it directly against users' money and data** — the critics are right, you're manufacturing risk. Same tool, two different outcomes. The dividing line is the word **"phase"**. ## Maturity isn't picking a side — it's knowing what phase you're in The real signal from this whole debate: **don't pick a side; distinguish the phase**. And "which phase am I in right now, which ruler should I use" — that judgment is exactly the product manager's job. AI has made *building* nearly free, which means **the judgment about doing the right thing at the right phase** has become the most valuable thing in the room. > High-fidelity first — but a prototype is never the same as shipping. Trust AI's speed; hold the production line. doaipm's five-phase workflow and the safety net are designed precisely for "using the right tool for the right phase." Start at the [method center](/en/method/) and the [言出法随 ("Speak it, AI builds it") playbook](/en/method/playbook/). --- **Further reading** - Vibe Coding — [The Vibe Coding Debate 2026: Both Sides, Sourced](https://vibecoding.app/blog/vibe-coding-debate) - Smashing Magazine — [When "Production-Ready" Becomes a Design Deliverable](https://www.smashingmagazine.com/2026/04/production-ready-becomes-design-deliverable-ux/) - Userpilot — [Product Management in 2026: Is AI Product Management a Lie?](https://userpilot.com/blog/what-is-product-management/) --- # Let AI execute, keep the judgment yourself: in 2026, the PM role is being redrawn URL: https://doaipm.com/en/blog/from-executor-to-orchestrator/ Published: 2026-06-02 Tags: doaipm, AI-native PM, agentic, 言出法随 In 2026, something quiet but thorough is happening to the product manager role: **it's being redrawn**. Not eliminated — the work itself has been cut differently. Some of it goes to AI. Some of it lands back on you, heavier than before. Start with a few numbers. Industry data puts the average PM at spending roughly **30%** of their time gathering and synthesizing information, **20%** on communication and alignment, and just **15%** on actual strategic thinking. What AI is doing is straightforward — it's collapsing that first block to near zero. ## From executor to orchestrator Today's AI agents can run multi-step workflows on their own: monitoring user signals, triaging feedback, proposing roadmap changes, even kicking off an A/B test. That means the PM role is shifting **from executor to orchestrator** — you're no longer doing every step yourself; you're directing, reviewing, and deciding. That's not bad news. Manually grooming requirements, pulling data by hand, chasing stakeholders for alignment — none of that was where the job's value lived anyway. Handing it off is a form of liberation. ## Where to invest the time you've won back So the obvious question: what do you do with the hours you've recovered? The PMs running ahead of the field in 2026 give remarkably similar answers: **reinvest the reclaimed time in exactly the places AI can't reach** — product vision, user empathy, judgment, and taste. AI is excellent at processing data, recognizing patterns, and generating content. What it can't give you is **strategic intuition, the human politics of stakeholders, ethical trade-offs, and the quiet instinct to see a need others haven't seen yet**. Put it another way: AI owns *how*. What it hands back to you — harder, and worth more — is **what, why, and whether we got it right**. ## The twist: this isn't about doing less — it's about building things yourself If you picture "orchestrator" as doing less with your hands, you're reading it backwards. A scene that's becoming increasingly common in 2026: **product managers building their own tools** — custom dashboards, roadmap visualizations, PRD chatbots, quick-and-dirty data endpoints. Things that used to require queuing up engineering time. Now one person, one sentence, and it's done. Thinking and making are being reunited in the same hands — and those hands belong to the PM. This is exactly what doaipm has always been about: **言出法随 ("Speak it, AI builds it"), built on the raw capability of Claude Code** — the tool has to be powerful enough that you can afford to describe rather than operate. And **not knowing how to code turns out to be an advantage**: you're not constrained by a mental estimate of how hard something is to implement. You just describe what users need, clearly. ## Why judgment is the scarcest resource now There's another signal that tends to get overlooked. Gartner projects that **by the end of 2027, more than 40% of agentic AI projects will be cancelled** — due to runaway costs, unclear value, and insufficient risk controls. That tells you something precise: when *building* becomes nearly free, **deciding whether something should be built at all becomes the rarest thing in the room**. An agent that can run autonomously will not tell you whether it should exist. That's the PM's job — and it's only getting weightier. ## Day-to-day: execution goes to Claude Code, judgment stays with you Compress the whole shift into one actionable sentence: **execution goes to Claude Code; judgment stays with you.** - Lead with a one-sentence spec — state *what* you want clearly, then let AI restate it and push back, until you're sure it understood. - Use **high fidelity** to build real, runnable things directly — not wireframes. That's the speed-to-learning that the AI era actually makes possible. - Use the [high-fidelity verification checklist](/en/method/playbook/) to confirm you got it right — real content, four states, real interactions, the critical path, adversarial inputs. The title is fragmenting (AI PM, API PM, Agent PM…), but the core of the product manager role — **thinking something through clearly, saying it clearly, and owning the outcome** — hasn't been diluted. AI has amplified it. > Execution can be outsourced. Judgment cannot. The stronger AI gets, the more your judgment is worth. If you want to actually put this "let AI execute, keep the judgment yourself" way of working into practice, start at the [method center](/en/method/) and the [言出法随 playbook](/en/method/playbook/). --- **Further reading** - Product School — [AI Product Managers Are the PMs That Matter in 2026](https://productschool.com/blog/artificial-intelligence/guide-ai-product-manager) - Userpilot — [6 Product Management Trends in 2026: The PM Role Is Splitting](https://userpilot.com/blog/product-management-trends/) - Paraform — [What Is an Agent PM? The New Role Startups Are Racing to Hire in 2026](https://www.paraform.com/blog/what-is-agent-pm-role-startups-hiring) --- # Stop Learning, Start Doing: The Only Thing Standing Between You and AI-Native PM Is Action URL: https://doaipm.com/en/blog/stop-learning-start-doing/ Published: 2026-06-02 Tags: doaipm, AI-native PM, 言出法随, Claude Code Someone asked me recently: "I want to transition into product management for the AI era — what should I study first?" My answer might surprise you: **Don't study anything. Just start.** ## Stop hoarding knowledge The old assumption was: before you do something, build up your knowledge base first. Learn the frameworks, learn the tools, learn the process — get "ready" before you touch anything. That habit doesn't work anymore. The reason is simple: **you will never know more than AI knows.** Whatever you spend three months grinding through, AI already has it loaded, on demand, available the moment you ask. In a world where knowledge is instantly accessible, spending your time *stockpiling* knowledge is the worst possible investment. The right move is the opposite: **start doing, and ask AI on the spot when you hit something you don't know.** Not "learn it, then do it" — but "do it, and learn as you go." Knowledge flows to you the moment you need it, instead of sitting in a warehouse collecting dust. ## The core of DO AI PM is "DO" This method is called **DO AI PM**. The heaviest word in that name is **DO** — doing. Not *learn* AI PM. Not *study* AI PM. **DO.** Getting started is the prerequisite for everything else. And **the core of DO is speaking**. In the AI era, "building something" is roughly equivalent to "describing clearly what you want." You say it. Claude Code builds it. Speak it, AI builds it — 言出法随. ## And "speaking" is the most basic skill a PM already has Here's the good news: **describing what you want, clearly, is already the product manager's core competency.** You don't need to write code. You don't need to understand architecture. There's no technical prerequisite. The one thing you have to do is articulate "what does the user need, and is the experience right?" — and you're already doing that every day. So **DO AI PM has no barrier to entry.** It doesn't care about your background, your education, or your technical level. If you can speak in plain terms, you can start. ## The only barrier is inaction If there's no knowledge barrier and no technical barrier, what's actually standing between you and becoming an AI-era product manager? Just one thing: **you haven't started yet.** Not "I'm not ready" — not "I need to study a bit more" — those are just polished ways of saying you're procrastinating. The real barrier is that you haven't opened Claude Code yet. Haven't said the first sentence. ## Open Claude Code today So stop bookmarking tutorials. Stop watching from the sidelines. Do one tiny thing right now: open Claude Code, say "help me build a…", and watch it build. That moment? You're already an AI-era product manager. Everything after that is just getting better at it. > Don't learn. Do. The core of DO AI PM is DO; the core of DO is SAY — and speaking, you already know how. Take your first step at the [method center](/en/method/) and the [言出法随 playbook](/en/method/playbook/). --- # Vibe coding is already obsolete — and that's great news for product managers URL: https://doaipm.com/en/blog/vibe-coding-is-product-management/ Published: 2026-06-01 Tags: doaipm, AI-native PM, vibe coding, Claude Code, methodology In 2026, Andrej Karpathy — the person who popularized the term "vibe coding" — got on stage and called it obsolete. What he proposed in its place he named **agentic engineering**, and he described it as a list of activities: writing design specs, supervising plans, inspecting diffs, writing tests, building evaluation loops, managing permissions, preserving quality. Read that list again with the engineer-specific words removed. Deciding *what* should be built. Checking whether the result matches the intent. Defining what "good enough" means and refusing to ship below it. That isn't a new discipline. As Jeff Gothelf put it, **the judgment version is the job — it was always the job.** AI just removed the places we used to hide it. That is the single most important shift for anyone building products right now, and it points in a direction most people find counterintuitive. ## The skill that survives isn't coding For a decade, "learn to code" was the universal advice for anyone with a product idea. In 2026 the advice quietly inverted. When AI agents can turn a clearly-stated problem into working software, **documentation stops being the job — intent, clarity, and judgment become the job.** The industry has a name for the people who matter now, and it's not "prompt engineer." Product School's read on the year was blunt: AI product managers are the PMs that matter in 2026. The mechanics back this up. Across the field, the consensus on what an AI-native PM actually does has converged on three verbs: - **Describe** what you want clearly enough that a capable agent can act on it. - **Decompose** the problem into pieces small enough to verify. - **Judge** whether the output matches the intent — and decide what happens at the 15% where it doesn't. None of those three are coding. All three are product management. ## Why not knowing how to code can be an advantage Here's the part that sounds wrong and isn't. People who can code carry a running estimate of *how hard this is to build* in the back of their minds. That estimate is useful when you're the one building — and a liability when you're deciding what *should* exist. The idea gets pruned before it's spoken, shaped by implementation cost instead of user value. If you can't code, you don't have that reflex. You're left holding the only questions that matter at the start: what does the user actually need, and is this experience right? Then your job is to say it clearly. And "saying the idea clearly" is the oldest core skill a product manager has. That's not an argument for staying ignorant. It's an argument that the bottleneck moved. The scarce resource is no longer typing speed in a text editor; it's the clarity of the thing you're asking for. ## "Speak it, AI builds it" is a method, not a vibe The catch is that *vibe* coding earned its obituary for a reason. When you prompt loosely, every run drifts: designers have noticed that AI prototypes fill the space wireframes used to occupy, but each run of the same prompt produces something subtly different — semantic drift that widens fast. Loose intent in, inconsistent product out. So the answer isn't to prompt harder. It's to work with discipline. That discipline is what we call **DO AI PM**, and it comes down to a few commitments: - **High fidelity first.** Skip the wireframe. Build the real, runnable thing — real content, real states (loading, empty, error, success), real interactions — and verify it by actually running it. A working prototype outperforms a described one every time, including in the room where the decision gets made. - **Five phases, small steps.** Discover → Define → Design → Develop → Validate. In *Define*, make the AI ask you questions before it writes the spec. In *Develop*, change one thing at a time. - **A safety net.** No real secrets or production data in a prototype. The human presses the irreversible buttons — publish, delete, pay. Don't touch production. When you're unsure, ask. That last list is the difference between "I got lucky with a prompt" and "I can do this again on Monday." ## The proof is the work A method that only explains itself is worth nothing. So everything here is downstream of real products, all built this way with Claude Code: [SoloMD](/en/products/) (a Markdown editor), Unterm (a terminal AI agents can drive), unfetch (a download manager for humans and AI), StoryAlter (an AI writing companion), and Unflick (a media player for humans and AI). This very website — eight languages, this blog, the decks — was spoken into existence the same way. Not a line of it was hand-written. > The method explains the work; the work proves the method. Karpathy was right that vibe coding is over. What replaces it isn't a harder kind of engineering — it's the discipline of deciding well and describing clearly. That has a name, and it's a good time to have the job. If you want to feel what "speak it, AI builds it" is actually like, start from the [method center](/en/method/). --- **Further reading** - Jeff Gothelf — [Karpathy said vibe coding is obsolete. What he described instead is product management.](https://jeffgothelf.com/blog/karpathy-said-vibe-coding-is-obsolete-what-he-described-instead-is-product-management/) - Product School — [AI Product Managers Are the PMs That Matter in 2026](https://productschool.com/blog/artificial-intelligence/guide-ai-product-manager) - monday.com — [Vibe coding for product managers: the complete 2026 implementation guide](https://monday.com/blog/vibe-coding/vibe-coding-product-managers/) --- # Speak it, AI builds it: I made this website with a single sentence URL: https://doaipm.com/en/blog/welcome/ Published: 2026-05-30 Tags: doaipm, Claude Code, methodology The website you're looking at — its multilingual structure, the blog system, the presentation decks — **not a single line of code was written by hand**. I just described what I wanted, sentence by sentence, and let Claude Code do the rest. That's the whole point of doaipm: **Speak it, and AI builds it.** ## Not knowing how to code is an advantage Many people think "I can't code" is a handicap for building products. My experience is the opposite. People who can code carry "how hard is this to implement" in their heads, and they limit themselves before the idea is even out. When you don't code, you don't have that baggage — you only care about **what the user needs and whether the experience is right**, and then you say it clearly. And "saying the idea clearly" is exactly the core skill of a product manager. ## How this site was "spoken" into being - I said "build a multilingual training site," and it scaffolded an 8-language structure; - I said "foreground doaipm, downplay the personal brand," and it rearranged the nav and footer; - I said "add sponsorship and contact," and it wired up the channels and QR codes; - I said "I want to write the blog in Markdown later," and the system you're reading was born. At every step, I never touched a technical detail. My only job was to **think clearly and speak plainly**. ## The work is the proof Talking about a method proves nothing. So this site also lines up a row of real products — SoloMD, Unterm, unfetch, StoryAlter, Unflick — all built with the same method, all with Claude Code. > The method explains the work; the work proves the method. If you want to feel what "speak it, AI builds it" is like, start from the [Method center](/en/method/).