
DeepSeek's Invisible Founder Just Talked for 4 Hours. It Leaked.
What this covers
DeepSeek is a household name. Liang Wenfeng, who founded it, has given two real interviews, ever. No podcasts, no keynotes, no essays. Then in May he got on a call with his investors, closed door, and talked for almost four hours. Last week a 34,000-word transcript of that call hit WeChat. The original was taken down within a day; copies are everywhere now.
He came to AI from a quant fund, and it shows: he talks about intelligence the way a retailer talks about goods. Cost, price, margin, discipline. Ten months to pay back a server, so the API earns "sixfold profit" and no more. Give the weights away, because the recipe isn't the restaurant. Refuse consumer retention, video generation, world models, overtime, everything off the main line, and spend the difference on one bet about continual learning, made with the whole company. This is what four hours leave behind, including the parts that argue against him.
If you like this kind of breakdown, subscribe, and tell me in the comments which part of the transcript you think ages worst.
## Chapters
00:00 An invisible founder, and a four-hour leak 00:57 Top exam score, quant fund, ten thousand GPUs 01:36 Why he talks about AI like a retailer 02:01 The pricing rule: ten-month payback, "sixfold profit" 02:49 The Costco playbook, run on tokens 03:09 Open weights: the recipe, not the restaurant 03:55 Sesame seeds and the watermelon — what DeepSeek refuses 04:51 No overtime, "because it's just not that hard" 05:15 The real bet: language, agents, continual learning 06:40 Two caveats — and Yann LeCun betting the other way 07:22 Half the company is labeling data by hand 07:47 Turn the cash into GPUs — and the compute gap 08:37 "Nvidia is digging its own grave" 09:35 The one thing he won't give up 10:10 What four hours leave behind
## Sources & further reading
- The leaked investor-meeting transcript (~34,000 words, meeting dated May 20, 2026) — https://x.com/FuckAnthropic/status/2080183070842638813 - 新浪创事记 / 盒饭财经 — "这份42页的PDF,确实很'梁文锋'" (analysis of the leak and why the voice matches) — https://finance.sina.com.cn/tech/csj/2026-07-23/doc-iniiuwtf5389545.shtml - CNBC — DeepSeek reportedly told investors: no poaching our people — https://www.cnbc.com/2026/06/18/no-poaching-our-people-chinas-deepseek-reportedly-tells-investors.html - 暗涌 Waves (36氪) — Liang Wenfeng's rare on-record interview: "中国的AI不可能永远跟随" — https://finance.sina.com.cn/tech/2025-01-26/doc-inehhksk9178057.shtml - 快科技 — Liang as Zhanjiang's 2002 gaokao top scorer, archival photos and reports — https://finance.sina.com.cn/tech/roll/2026-06-10/doc-iniaxkrn5909653.shtml - New York Magazine — D. E. Shaw, the first great quant hedge fund (where Bezos worked) — https://nymag.com/intelligencer/2018/01/d-e-shaw-the-first-great-quant-hedge-fund.html - The Seattle Times — full Q&A with Costco founder Jim Sinegal on the markup cap — https://www.seattletimes.com/business/complete-qa-with-costco-founder-ceo-jim-sinegal/ - Columbia Engineering — Yann LeCun on why language models alone won't reach human-level intelligence — https://www.engineering.columbia.edu/about/news/metas-yann-lecun-asks-how-ais-will-match-and-exceed-human-level-intelligence - Andrej Karpathy — Software 2.0 (the dataset is the source code) — https://karpathy.medium.com/software-2-0-a64152b37c35 - DeepSeek-V3 technical report (hand-tuned Nvidia PTX) — https://arxiv.org/html/2412.19437v2 - DeepSeek — V3.2 release, kernels shipped in both CUDA and TileLang — https://x.com/deepseek_ai/status/1972604780469239843 - DeepSeek — V4 announcement — https://x.com/deepseek_ai/status/2047516922263285776 - Reuters — Huawei's Ascend supernode to support DeepSeek V4 — https://whbl.com/2026/04/23/huawei-ascend-supernode-to-support-deepseek-v4/ - deepseek-ai/TileKernels (still requires CUDA today) — https://github.com/deepseek-ai/TileKernels
Source description (no synthesized summary yet).
Liang Wen, DeepSeek's founder, operates an AI company with a quant fund's discipline: ruthlessly optimizing for cost and capability by capping margins, releasing open-source weights, refusing non-core products, and concentrating all resources on advancing toward AGI through language models and continual learning.
- DeepSeek prices APIs with 6x hardware cost margins (10-month payback on 3-5 year lifespan) and deliberately undercuts demand to prioritize adoption over revenue
- Liang believes human intelligence is fundamentally linguistic, so language models are the seed of AGI, with the next frontier being continual learning rather than multimodality or other specializations
- The company ruthlessly refuses peripheral products (video generation, world models, consumer apps) to concentrate compute and talent on the main line to AGI, even as millions of users arrive without retention effort
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
High Flyer bought GPUs aggressively before their purpose for AI was widely known: 100 cards in 2015, 1,000 in 2019, 10,000 by 2021. Liang says this purchasing was driven by curiosity.
“Along the way, the fund kept buying GPUs long before anyone knew what for. 100 cards in 2015, 1,000 in 2019, 10,000 by 2021. He says it was curiosity.”
In the 'Software 2.0' framework (coined by Andrej Karpathy), the dataset is the source code and the people labeling data are the programmers; this reframes data labeling from a cost center into a core engineering activity.
“Karpathy coined the frame for this years ago. Software 2.0. The data set is the source code, and the people labeling it are the programmers.”
Half of DeepSeek's workforce, including half the core researchers, spends time labeling data by hand; high-end data labeling costs the same in China as in the United States, so there is no cheap labor cost advantage.
“Meanwhile, the unglamorous present. Half of DeepSeek, he says, including half the core researchers, is labeling data by hand. High-end labeling costs the same in China as in America. There's no cheap labor shortcut.”
Liang concludes that the gap between Chinese and US AI capability is fundamentally a compute gap, and because fewer chips means fewer experiments, the talent gap is downstream of the compute gap.
“His conclusion, the gap with the US is compute, and even the talent gap is downstream of it, because fewer chips means fewer experiments.”
DeepSeek is deliberately text-first and not pursuing multimodality as a core research area because multimodality is a useful component tool (like search) but is not intelligence itself and does not contribute directly to AGI.
“Even vision. DeepSeek is famously text-first, and his reason is that multimodality is a component, useful, though a search is useful. It isn't intelligence itself.”
No one currently has a working method for continual learning; inside DeepSeek, this type of exploratory research is called 'Mo Jang' (feeling for the prize), and researchers describe it as lottery tickets: cheap to try, anyone can think about it, but nobody knows in advance which approaches will succeed.
“Two honest caveats belong here. First, nobody has a working method for continual learning. He says so himself. Inside DeepSeek, this kind of research is called Mo Jang, feeling for the prize. Lottery tickets. Cheap to try, anyone can think about it, nobody knows who wins.”
Serious figures like Yann LeCun argue that language models alone cannot reach human-level intelligence; they contend that world models—systems that can predict the effects of actions—are necessary for AGI. Liang is betting the entire company against this thesis.
“Second, serious people bet the other way. Yann LeCun argues language models alone won't reach human-level intelligence. You need world models that predict what actions do. Liang is betting against that with the whole company.”
A transcript of a closed-door investor call with Liang Wen lasting almost four hours and totaling 34,000 words leaked to WeChat in May; the original was removed within a day, but copies are now ubiquitous; the meeting was independently corroborated and the voice matches Liang's two public interviews, establishing authenticity.
“Then in May, he got on a call with his investors, closed door, and talked for almost four hours about everything. This week, a 34,000 word transcript of that call hit WeChat. The original was taken down within a day. Copies are everywhere now. DeepSeek hasn't confirmed it. That's how leaks work. But the meeting was independently corroborated, and the voice matches those two interviews exactly.”
DeepSeek has lost key researchers to rivals, and the funding round reportedly came with a binding condition: no poaching of DeepSeek staff and no funding of DeepSeek spinouts, suggesting that retention is a competitive battleground.
“It's backed by behavior. They've lost key researchers to rivals, and the funding round reportedly came with one condition for every investor. No poaching Deep Seek staff, no funding spinouts.”
Liang Wen is not a household name despite DeepSeek's prominence; he has given only two formal interviews ever (no podcasts, no keynotes, no essays), making him one of the least publicly visible major AI figures.
“Liang is not. He has given two real interviews ever. No podcasts, no keynotes, no essays. The man behind the company is basically invisible.”
The primary job of the next model is to be useful to DeepSeek itself—to help DeepSeek's researchers build the subsequent model—creating a feedback loop where AI accelerates AI research, which is why Liang expects progress to compound rather than arrive as a single step change.
“And the next model's first job, in his telling, is to be useful to DeepSeek itself. A model that helps its own researchers build the one after it. AI accelerating AI research, which is why he expects progress to compound, not arrive in a single moment.”
In Liang's framing, apps (like consumer interfaces) and APIs are 'sesame seeds' while AGI is the 'watermelon'; the company picks up seeds on the way but doesn't stop to harvest them, meaning peripheral products are accepted passively but not pursued strategically.
“In his metaphor, the app and the API are sesame seeds. AGI is the watermelon. You pick up seeds on your way, you don't stop for them.”
The speaker's overall interpretation is that Liang operates with 'a quant's discipline applied to intelligence itself': cap margins, cut costs, open the recipe, refuse everything off the main line, and spend the savings on one bet (AGI via language models and continual learning).
“It's the picture Four Hours leave behind. A quant's discipline applied to intelligence itself. Cap the margin, cut the cost, open the recipe, refuse everything off the main line, and spend the difference on one bet.”
Liang's foundational hypothesis is that human intelligence might be fundamentally linguistic—that thinking is the brain weaving words together—and therefore a language model is the seed of general artificial intelligence.
“Now, the main line itself. His foundational bet goes back to that 2023 interview. Human intelligence might fundamentally be language. Thinking might be the brain weaving words. If that's true, a language model is the seed of general intelligence.”
DeepSeek currently operates at 20,000 H100 equivalent compute (the available Nvidia chip in China), while the largest American frontier models run approximately 800 billion active parameters against DeepSeek's 49 billion.
“He puts DeepSeek at 20,000 H equivalents, the China market Nvidia chip, while the biggest American models run around 800 billion active parameters against DeepSeek's 49.”
Despite DeepSeek's compute constraints, Liang argues that Nvidia is 'digging its own grave' because CUDA's moat—the software lock-in preventing developers from rewriting code—is eroding as AI writes code faster than humans, making CUDA rewriting feasible.
“And yet, in the same meeting, Nvidia, he says, is digging its own grave. CUDA's moat is software nobody wants to rewrite, but AI writes code now.”
DeepSeek's pricing strategy parallels Costco's approach of deliberately capping markups at about 14% for decades; the low-margin cap creates trust and makes it impossible for competitors to undercut while being profitable themselves, forcing rivals into territory DeepSeek has already made its home.
“This is a known playbook. Costco has capped its markup at about 14% for decades, and the cap is exactly why people trust the store. Keep margins low on purpose and anyone who wants to undercut you has to live in territory you've already made your home. Deep Seek is running that playbook on AI tokens.”
Liang's career path from quantitative finance to AI mirrors Jeff Bezos's path from D.E. Shaw to Amazon, and this background explains DeepSeek's focus on cost, price, margin, and operational discipline.
“That path should sound familiar. Jeff Bezos spent four years at a quant fund, D.E. Shaw, before leaving to start Amazon. Liang went from a quant fund to AI, and it shows. In this transcript, he talks about AI the way a retailer talks about goods. Cost, price, margin, discipline. Everything DeepSeek does makes more sense heard in that language.”
DeepSeek deliberately launched one model at a high price to constrain demand due to fears of overwhelming the system; when price was cut to one-quarter, internal company chat 'erupted in cheers,' revealing that the team prioritized user adoption and ecosystem growth over revenue.
“One of their models launched with a high price on purpose. They were worried about too much demand. The team hated it. He cut the price to a quarter and the company chat erupted in cheers. The point of the model for them was people using it, not the revenue line.”
DeepSeek's open-source strategy does not kill the business because model weights are the recipe (replicable) while serving a frontier model cheaply is the hard part (infrastructure, kernels, cluster optimization, utilization); a competitor can download DeepSeek's open model tomorrow but would still need to serve it at multiples of DeepSeek's cost.
“Giving away the weights doesn't kill the business, he argues, because the weights are the recipe, not the restaurant. Serving a frontier model cheaply is the hard part. The kernels, the clusters, the utilization. A rival can download the model tomorrow and still serve it at multiples of his cost.”
DeepSeek avoids overtime and brutal work schedules because research requires relaxed minds, and the company refuses so much non-core work that there simply isn't enough work to create overtime pressure; products ship imperfect and remain imperfect rather than being polished through extended effort.
“One more consequence of refusing work, DeepSeek doesn't do overtime in an industry famous for brutal hours. His explanation? Research needs relaxed people, and they refuse so much work that there isn't that much work. Products ship imperfect and stay imperfect. Quote, "We don't need to work overtime because it's just not that hard."”
The critical gap in current AI is continual learning: a new employee learns company culture, relationships, and processes over two months, then can execute an instruction like 'Go get Xiao Wang' with full contextual understanding; in contrast, an AI system requires all context spelled out explicitly every time, and the next model must learn continuously the way the employee learned.
“His example of what's missing is concrete. A new employee spends two months learning the company, who's who, how things work. Then you say, "Go get Xiao Wang." and they just know. An AI needs all of that spelled out every time. Who Xiao Wang is, where he sits, how to approach him. Liang's claim, with complete context, AI already beats the human. But you can never supply complete context. So, the next model has to learn continuously, the way the employee did.”
Every card Nvidia sells trains better AI, which ports the ecosystem faster—creating a positive feedback loop where Nvidia's customer base accelerates the development of tools that make Nvidia obsolete.
“Every card Nvidia sells trains better AI, and better AI ports the ecosystem faster.”
Inside DeepSeek, there are no KPIs (key performance indicators), and engineers receive half their time for self-directed research, enabling exploration and preventing the metric-driven culture that often stifles innovation.
“Inside, no KPIs, and engineers get half their time for self-directed research.”
DeepSeek's pricing rule is to set API prices such that hardware costs are recovered in 10 months; since servers run 3 to 5 years, this yields approximately 6x the hardware cost as profit over the server's lifetime, which Liang explicitly describes as 'we only earn sixfold profit.'
“Start with the pricing rule because it's one sentence. Buy a batch of servers, price the API so they pay for themselves in 10 months. Servers run 3 to 5 years, so 10-month payback works out to roughly six times the hardware cost over a server's life. His words, we only earn sixfold profit.”
DeepSeek built TileLang, a high-level language that rewrites GPU kernels with a fraction of the manual effort required, and uses AI to write TileLang itself, demonstrating practical ecosystem escape from CUDA dependency.
“DeepSeek built TileLang, a high-level language that rewrites GPU kernels in a fraction of the effort, and uses AI to write TileLang itself.”
The open-source model DeepSeek releases is the same as the deployed model: identical weights, not a weaker public release.
“He was also direct about the version question. The open model is the deployed model. Same weights, not a weaker public release.”
The speaker's concluding promise is that the strongest model DeepSeek builds will always be released open-source, identical to the version DeepSeek runs internally.
“It also comes with a promise you can check. He said the strongest model they build will always be released open, the same one they run themselves.”
When asked what is non-negotiable for DeepSeek's success, Liang identified the team (not compute or models), saying 'If the team stays, we achieve AGI. Everything else is just time.'
“Asked for his one non-negotiable, he didn't say compute or the models, he said the team. If the team stays, we achieve AGI. Everything else is just time.”
DeepSeek's escape from CUDA is unfinished: V3 trained on Nvidia with hand-tuned Nvidia instructions; V3.2 shipped kernels in both CUDA and TileLang; V4 achieved day-zero support on Huawei's Ascend chips, but DeepSeek's Tile kernels library still requires CUDA as a backend today.
“V3 trained on Nvidia, down to hand-tuned Nvidia instructions. V3.2 shipped kernels in both CUDA and TileLang. V4 got day zero support on Huawei's Ascend chips, and DeepSeek's own Tile kernels library still requires CUDA today.”
When asked whether human taste still matters in AI development, Liang said AI already has taste and intuition; what it lacks is continual learning.
“Asked whether human taste still matters, he said AI already has taste and intuition. What it lacks is continual learning.”
When R1 went viral and millions of consumer users arrived overnight, DeepSeek made almost no effort to retain or monetize them; the API runs with minimal staff (only a few maintainers, no sales team, no customer service), reflecting Liang's view that consumer apps are distractions from AGI research.
“When R1 went viral last year, millions of consumer users showed up overnight, and the company did almost nothing to keep them. No retention push, no monetization. His words, "We really couldn't chase them away." The API runs with a few maintainers, no sales team, no customer service.”
Liang wants thousands of companies built on top of DeepSeek; each company shapes the ecosystem to be more DeepSeek-dependent and increases switching costs and network effects.
“And he says he wants thousands of companies built on top of Deep Seek. Each one makes the ecosystem a little more DeepSeek-shaped.”
Open-source models only threaten a company if that company is charging 100x the cost of the open alternative; at DeepSeek's pricing, running the business model of DeepSeek using DeepSeek's own open weights would lose money for competitors.
“At Deep Seek's price, competing with Deep Seek using Deep Seek's own model loses money. Open source only threatens you, he says, if you're charging a hundredfold.”
DeepSeek refuses to build video generation capabilities because, despite being 'a good business,' video generation is not on the path to AGI and would divert resources.
“Same answer for whole product categories. Video generation, "A good business," he says, "and not on the road to AGI. So, no."”
Robotics is absent from Liang's AI research staircase because robotics is labor-intensive; Liang's strategy is to solve continual learning and other core problems first, then deploy AI to handle the 'bitter work' of robotics rather than investing in robotics now.
“And robotics sits lost on his staircase precisely because it's labor-intensive. Solve learning first, then let AI do the bitter work.”
Liang plans to deploy billions of raised capital into purchasing GPUs (specifically H100 equivalents, Nvidia's China-available chips) as quickly as possible; his reasoning is that cash in the bank earns 2% while a GPU pays back in 10 months, so the opportunity cost of not spending immediately is significant.
“Back to the businessman lens, applied to hardware. The plan for the billions raised, turn them into GPUs, as many as possible, as fast as possible, a premium is fine. His reasoning, cash in the bank earns 2%. A GPU pays back in 10 months. If I could spend it all in 6 months, he says, that would be wonderful. The urgency is real.”
DeepSeek refuses to invest heavily in world models as a core product direction because world models are not currently the bottleneck to AGI.
“World models? Not the bottleneck right now.”
Liang's proposed staircase to AGI is: chain of thought (completed), then agents (current year), then continual learning (next frontier). Continual learning is the critical missing capability.
“From there, his staircase, chain of thought last year, agents this year, and the next step, continual learning.”
Liang founded High Flyer, a quant fund that trades markets with machine learning, and built it into one of China's largest quant funds.
“Then he founded High Flyer, a quant fund that trades the markets with machine learning, and built it into one of China's biggest.”
Liang predicts that within one year, the domestic (Chinese) chip software problem will be solved, and the remaining bottleneck will shift to Huawei's production capacity for Ascend chips.
“His prediction, within a year, the domestic chip software problem is solved, and the bottleneck that remains is Huawei's production capacity.”
To match the scale of American frontier models, training at DeepSeek's current scale would require approximately 50,000 next-generation chips.
“Training at that scale would take, by his estimate, 50,000 next-gen chips.”
When someone on the investor call challenged the 6x profit margin as too high, Liang agreed there was room to cut but noted that at the current price, demand no longer increases, so he leaves it unchanged.
“Someone on the call pushed back. That's too profitable. He agreed and said there's room to cut, but at this price demand no longer moves, so he leaves it.”
DeepSeek's loose organizational structure is currently straining, and the company is actively building real hierarchy now, suggesting that the early-stage flat/exploratory culture is becoming unsustainable at current scale.
“He also admits the loose structure is straining. Real hierarchy is being built right now.”
Liang Wen achieved the top score on the college entrance exam in Zhanjiang in 2002 and attended Zhejiang University for engineering.
“In 2002, he got the top score on the college entrance exam in his home city of Zhanjiang, and went to Zhejiang University for engineering.”
DeepSeek's research team was spun out from High Flyer in 2023.
“In 2023, the research team became DeepSeek.”