7/26/2026 at 12:19:08 AM
I think the way to parse the current title "DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]" is that there was a leak that DeepSeek will pause fundraising because they perceive there is a compute gap with the US.I am also guessing that the majority of the people who read this title will think that DeepSeek is pausing this fundraising because some comments they made about the compute gap were leaked. That is not the case.
by credit_guy
7/26/2026 at 12:43:32 AM
Maybe: "Leaked Deepseek transcripts reveal plan to pause fundraising due to compute gap"I don't know what "compute gap" means in this context though and it's not clear that that's why they plan to pause fundraising or if the title is conflating.
by culi
7/26/2026 at 10:01:39 AM
yep. the word they use is probably 克制 or self-restraint. no need to raise so much cash if you can't use it.in his article he talks about the negative aspects of getting everything you want. (all the money, brightest minds, biggest share in AI) etc. he says that these are the things that will cause a company to fail.
by choonway
7/26/2026 at 10:59:55 AM
I shouldn't win too hard, because then I'll lose?by andai
7/26/2026 at 11:14:27 AM
Yes. Once you're in such a comfortable position that you'll keep making lots of cash regardless of whether you do well or not, there is little incentive to make good decisions, let alone take risks. This is known as the "curse of Oil" and is also what Intel's decline is attributed to.by muvlon
7/26/2026 at 12:41:55 PM
Catching up without overtakingby selimthegrim
7/26/2026 at 1:50:35 PM
Sudden availability of capital and the perceived need to be seen doing something with it can be a curse. WeWork comes to mind, eg. Stay lean and mean until you actually need the capital. A company like DeepSeek will have zero issues raising anytime.by c7b
7/26/2026 at 1:19:09 PM
Yep. Classic “giant series A, company gets amazing office space” vibes.by brookst
7/26/2026 at 3:24:13 PM
> I shouldn't win too hard, because then I'll lose?Yes. It's a known thing.
by lelanthran
7/26/2026 at 8:41:39 PM
Raising more money ≠ winning harder.by DANmode
7/26/2026 at 3:38:10 PM
“Raising cash” isn’t necessarily “winning” since you have to give up equity/control, usually.by lotsofpulp
7/26/2026 at 1:19:39 PM
Problem of the local maxim, extreme dependency is the same as evolutionary pressure to specialize for it, eventually the process owns you.by Psyladine
7/26/2026 at 4:08:37 PM
Unless they have a shortage of mathematicians, physicists, ... it seems a fundraiser could help bypass a compute gap by focusing even more on inference and training efficiency.by DoctorOetker
7/26/2026 at 11:07:57 AM
is this similar to Meta starting to rent out own compute as they cant seem to do much with it and monetising it is much better ROI?...by antiford2049
7/26/2026 at 11:28:29 AM
Sorry to barge in here. I couldn't find a good place to placehttps://github.com/demo-zexuan/liang-wenfeng-investor-meetin...
The current link is 404, can mods update to above, detach, make sticky?
(No response from mods, understandable)
by oliculipolicula
7/27/2026 at 2:15:45 AM
I'm not sure I understand the question but I've replaced the top link (https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...), which was 404ing, with the link in your comment here. Does that help?by dang
7/27/2026 at 2:46:54 AM
Yes, thank you!It's now dropped off the FrontPage but I suppose I will just repost at some point with a less controversial title
by oliculipolicula
7/26/2026 at 11:40:13 AM
email hn@ycombinator.comby fragmede
7/26/2026 at 7:29:30 PM
The maturity of leadership in China seems to be on a whole different level from the U.S.Lots of companies in the U.S. have fallen victim to that syndrome, but if you used the word "restraint" in that context in Silicon Valley most people would look at you like you're insane.
by koverstreet
7/26/2026 at 2:10:53 AM
Yeah, wouldn't it makes sense to increase fundraising, so as to acquire more compute to close the gap?by AnnikaL
7/26/2026 at 10:09:22 AM
I guess they likely can't right now due to no hardware available, either cause of bans or already all booked.by epolanski
7/26/2026 at 11:25:08 AM
I gather this is the whole point of the original article. It seems to have been pulled, so I can't confirm.At what point do the AI companies start building their own hardware, I wonder.
by aae42
7/26/2026 at 12:43:40 PM
Other companies are investing and hardware design and manufacturing, so they won't duplicate the effort given the national drive it would be a waste.by hirako2000
7/26/2026 at 7:57:03 AM
Perhaps it's akin to a mineshaft gap?by GolfPopper
7/26/2026 at 2:08:33 AM
> if the title is conflatingThe title is certainly a great conflation. Any seeker of capital would want to regroup after an unfiltered leak of this magnitude, if for no other reason than to secure the forum from future leaks. The comments about the unlikelihood of enormous future profits were at least as consequential with regard to capital investment as anything else that was said.
by topspin
7/26/2026 at 2:07:47 AM
I skimmed the doc and my impression is that your second listed interpretation -- DeepSeek is pausing investment because of a leak -- is the more correct one.There's quite a bit of confidential information in the doc about the company and how it's positioning itself going forward to compete with US labs. I'd imagine they're not happy at all with this being leaked and are withholding investment as a punitive measure.
Not to mention the other interpretation seems illogical -- why would you pause fundraising if your perception was that you lacked resources compared to your competitors?
by ansk
7/26/2026 at 11:38:22 AM
Because money isn’t free and if you know there is a chokehold in supply, why raise money ar current valuation when things are getting better by the day?by asdf88990
7/26/2026 at 9:08:26 AM
Maybe because if you need more money to catch up with your competitors than anticipated then your ROIC is lower and your valuation changes?by tedggh
7/26/2026 at 12:23:03 PM
[dead]by z0ltan
7/26/2026 at 2:47:15 AM
All the Chinese reporting I see point to the second (majority) interpretation. Liang being furious about his private investor talk leaked online is the news here.by crazylogger
7/26/2026 at 4:00:50 AM
Those could be subsequent developments, but that's not what the linked transcript was about. The transcript was a discussion of the DeepSeek founder (Liang Wenfeng) with investors, and he does not mention any leaks, or any frustration. He simply says that he is constrained by the supply of cards, and he has no problem of getting funding, but has no reason to raise further funding because he can't transform the cash into cards. > There is certainly no shortage of funds or resources --- in fact, all these are readily available [...]
> Within our financial capacity, it's undoubtedly true that the more cards are always better. Our current strategy is to purchase as many cards as possible at a reasonable price --- exactly how many we can afford after using this funding round. The spending pace isn't predetermined; we'll buy whatever is available as long as prices remain competitive. In fact, I'd consider that a positive outcome if we spend the entire amount within six months. [...]
> In reality, spending such a large sum is no easy task: you can't obtain enough cards, they're hard to come by [...]
> Therefore, our only concern is whether we can obtain enough cards. If converting all funds into cards were feasible, we would undoubtedly do so without hesitation and are even willing to pay a premium for this benefit --- it's simply to cost effective. Even after paying the premium, however, achieving this goal remains challenging.
by credit_guy
7/26/2026 at 5:19:19 AM
he's pausing because of the leak. the transcript content has nothing to do with it.by try-working
7/26/2026 at 12:21:37 AM
Thank you! That is indeed how I read it.by solarkraft
7/26/2026 at 12:55:20 AM
I wanted to post this which explains the wording but I thought the transcript was more interesting. Sorry. Maybe mods can help me to put what follows as auxiliary link. I don't know how.https://www.bloomberg.com/news/articles/2026-07-25/deepseek-...
Update:
Less-paywalled word-for-word copy it seems at
https://fortune.com/2026/07/25/deepseek-liang-wenfeng-backer...
by oliculipolicula
7/26/2026 at 1:19:09 AM
Most of that is paywalled, but this one paragraph in the Bloomberg article suggests it might be more to do with investors leaking information:"The suspension stemmed in part from Liang’s frustration over online reports about his comments to investors during his first financing deal"
The part of the transcript I'd seen floating around online was this part from around 1 hour 26 min:
"With the largest models available today, we simply cannot afford to train them. Even if we spent all five hundred billion yuan, we still wouldn't be able to do so. Even if we could accumulate the resources, we wouldn't have the means to utilize them. The current largest model requires approximately 800 billion activations; domestically, we are still at a scale of several dozen billion activations, and even the largest domestic model may only require several dozen billion activations—a difference of an order of magnitude. To train a model of the same size as an AI system, we would need around 50,000 GB300 GPUs or Huawei 950 GPUs, totaling two hundred thousand cards. This is merely training; research has not yet been considered. Therefore, the biggest gap between us and the United States lies in resources."
by SyneRyder
7/26/2026 at 2:21:38 AM
amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal fast attention kernel yesterday, will be standing up cuda/metal/armv8 kernels too and thats gonna be fun.i genuinely think these models should be like 0.1 percent sparse for same capabilities we associate with them today, but theres no sane way to do that with extent tools. i built the right core tech for that in 2014 when there wasnt a market, but now there is and the experimentation velocity is wild.
amusingly llms really have a hard time using my simple apis because its not in distribution array programs. but i literally stood up cpu custom memory format and micro kernel for dense causal attention in less than 24-36 hours and outperforms the equivalent fused ggml/llama cpp fast oath by like 20-25 percent
by carterschonwald
7/26/2026 at 10:40:13 AM
I am not much of a math person, but if we look at the how the brain is wired, we see that the dendrites (the inputs) of a neuron are hundreds of micrometers in length, and the axons (the outputs) are millimeters, and very rarely can stretch to tens of centimeters. So they can sample only a tiny amount of internal state, and affect a much larger, but usually still small output.In math terms, this means a layer of a network can be represented with a block matrix in the whole 'layer' matrix, which I think means its sparse as you said.
As I said, my math knowledge is rusty, but I remember that a lot of matrix optimization techniques center around decomposing large matrices into these smaller blocks, which are then evaluated, and the output is combined in a final pass. Which leads to a huge reduction on parameter numbers and the time it takes to evaluate the result
by torginus
7/26/2026 at 3:24:25 PM
exactly. you certainly know more about the brain than i :)by carterschonwald
7/26/2026 at 3:17:07 AM
Look forwarding your future releasesby jacktang
7/26/2026 at 3:31:50 AM
i definitely will be doing some drop of some faster attention kernels in the next few weeks.like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen. so i can optimize the kernel flops
by carterschonwald
7/26/2026 at 6:56:02 AM
Curiosity:For most of the past five years, I've known ways to do better than Anthropic, OpenAI, and friends in many ways, at least on paper. I know I was right about many of them since many would show up 6-24 months later tools from the major providers, or otherwise become standard practice.
A central problem is the Mythical Man-Month. True, I could do those, beating then-state-of-the-art, but only given 2-5 years. I suspect many other people knew about them too and could do so as well. As I noted above, throwing people and dollars caused many of those to be built in less time than I could have regardless.
Other methods, I'm less confident about (>50%, <80%), but would lead to similar improvements orders-of-magnitude as you're predicting, but mine would need $$$$$ in compute and engineering infrastructure to build out. E.g. they need to not just theoretically work, but to try, I would need to convince someone to invest in them working.
So the TL;DR is that my knowledge was not at all helpful towards e.g. competing with OpenAI, Anthropic, or even building a small business.
However, where it was useful was in predicting where the industry was going. This is true in investing (but not easily, at least with my skill set), but in developing startups and systems, there were capabilities which I (correctly) assumed would be there, whereas there were many arguments that "AI will never be able to ____."
If I know how to do something, it will almost certainly happen, regardless of whether I'm the one who does it.
To be clear, my expertise is almost certainly nowhere as deep as yours. I'm not providing a direct analogy, or claiming others know what you do or can do the same. My point was really that if you believe you can have these models be 0.1 percent sparse for same capabilities we associate with them today:
a) You're probably right. They were built quickly for capabilities. A slower process can almost certainly lead to much smaller models too. That's a radical statement: Historically people claiming a 1000x improvement somewhere were crackpots, but that's very possible in an industry as fast-changing as this one.
b) Someone at Anthropic or OpenAI might be working on building out extent tools right now. Even if so, there are indirect ways to capitalize on that knowledge.
c) Critically, that predicts a future where Fable is $1/month instead of $100/month, and that's something which CAN be acted upon in planning.
It also suggests -- much less strongly -- the existence of much more sophisticated models at $100/month. There are open discussion in planning about whether models plateau, continue improving, singularity, or otherwise. That changes the biases there.
by blagie
7/26/2026 at 8:29:20 AM
>Fable is $1/month instead of $100/monthwill Anthropic (or OpenAI) lowers their price, or increase margin (to justify valuation)
by homarp
7/26/2026 at 10:56:37 AM
thx for the kind response!at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the performance claims at least on the dense side are solid so far!
the sparsity angle is because i want magic in the world. like anyone with a really chunky computer like any of those mac mini pros or serious workstation / server tier compute should be able to train from scratch their one 31b equivalent model in a week or so tops is the goal post i have in mind
by carterschonwald