7/21/2026 at 3:52:39 PM
I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones.Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.
edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.
edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.
by postalcoder
7/21/2026 at 4:15:16 PM
It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.by Tenoke
7/21/2026 at 9:27:10 PM
That's the only explanation that makes sense. If it was frontier but cost or compute were limiting factors, they'd release it at an obscene price for the bragging rights. Google doesn't care that much about alignment, and I don't think it's likely to be significantly different than 3.5 anyway. The only reason it would need to be soft-canceled is if it's terrible, and has to end up in a ditch like Llama 4 to avoid shareholder panic.by Miraste
7/21/2026 at 10:21:34 PM
3.5 pro was clearly a miss. It should have been in prod mid may, not MIA in late July. The brain drain at deep mind is a clear indicator that the people who know the most think that they can’t stay at the frontier.Antigravity NEEDED to be game-changing. Without the stream of data that Claude, Codex, and Cursor enjoy there is little chance of getting an effective reinforcement learning loop. For the first time in its history, GOOG is at a meaningful data disadvantage, and apparently a cultural one as well.
by reilly3000
7/21/2026 at 5:57:01 PM
That and/or the business case isn’t as clear when serving enormous models? You’re constantly stuck in a red queen’s race where your profitability window is increasingly measured in weeks because the Chinese are right behind you.For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.
by janalsncm
7/21/2026 at 10:06:12 PM
> For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See https://openrouter.ai/deepseek/deepseek-v4-flash#providers The developer antirez of Redis fame uses this as a local coding model.
DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.
So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.
by ekidd
7/21/2026 at 8:07:48 PM
The Chinese have been right behind OpenAI and Anthropic for ~18 months now.DeepSeek didn't do to OpenAI and Anthropic what nearly everybody claimed they would.
Every single person on HN that loudly proclaimed the end was nigh for GPT & Co. due to DeepSeek, was wrong. They were humiliatingly wrong, and they'll never own up to it. The reason those people were so very wrong, is the same exact reason the Kimi crowd is wrong now. And it's very obvious that they're wrong, but they have intense emotional blinders on. Their thinking process is hyper emotionalism: they want a certain outcome, regardless of if reality aligns to that or not. They're making emotional wishes about how they want things to turn out, and pretending those magic wishes are grounded in reason.
It takes enormous resources to run something equivalent to GPT 5.6 or Fable. Nobody can or wants to do that outside of very limited situations - if you can just reasonably pay as you go instead. As it turns out, you can just pay as you go with GPT and Fable. Their businesses have gotten radically larger since DeepSeek launched. Get it yet?
Domestic China is the only very large audience for their own models, so long as OpenAI and Anthropic stay top tier.
All the hype online from the forums about Kimi, is worthless: those people hyping it can't even come close to running it locally, which is the fantasy. So why are they hyping it? Why did they hype DeepSeek just the same, and learn nothing from its total failure to actually take down OpenAI and Anthropic? Rhetorical questions with obvious answers.
Kimi poses zero actual threat to OpenAI and Anthropic. Those companies will continue to pile up the subscriptions and API usage. Check out GPT's subscriber base today vs when DeepSeek launched. Get it yet? When Model X launches out of China in a year, we'll have this same conversations all over again, and the hypsters will have learned nothing.
While the Kimi fawning is endless, OpenAI will just keep piling up subscriber counts, and Anthropic will keep piling up API usage. Then OpenAI is going to staple a gigantic ad system onto GPT. China can't compete in the model-as-a-service business globally, for the exact same reason they failed so miserably to compete in search globally.
by adventured
7/21/2026 at 8:27:23 PM
> Domestic China is the only very large audience for their own modelsI don't think so. US models are very expensive, and not available in every country. I am not willing to pay $50/1M tokens for writing my pet projects.
by codedokode
7/21/2026 at 9:22:16 PM
There are also US based companies like Fireworks serving up the best open weight models with the compliances we need in US enterprise. Depending on the company, they may offer more/different jurisdictions, EU probably needs a Fireworks like company (haven't heard about one, maybe it already exists?)by verdverm
7/21/2026 at 9:48:16 PM
There is at least doubleword.ai, and there should be others.by crown421
7/21/2026 at 8:23:33 PM
without any hard data one way or another your comment is worthless. "pile up subscriptions" - based on what? neither company is public. "piling up subscriber counts", "piling up API usage"? cool. how much money are they making? oh you don't know because they're not public.the reality is one way or another that as long as there exists an alternative that a USA company could serve with the same compute rented from hyperscalers, this represents a threat, even if the extent to which is unknown
by amazingamazing
7/21/2026 at 8:24:49 PM
But what does that mean for Google if their model isn't as good as OpenAI's and Anthropic's?by Wowfunhappy
7/21/2026 at 9:00:58 PM
Why wouldn't the hyperscalers run these open models since they're much better than OpenAI and Anthropic at operating compute at scale?I think the reason OpenAI and Anthropic stay ahead in revenues right now is because the models are improving too quickly to reliably compete with them on cost.
But, once model performance reaches a plateau -- they have to at some point, though perhaps years away -- that's when ability to operate compute infrastructure at scale becomes the secret sauce.
The big AI labs are likely safe until models stop improving fast enough to protect them from competition on cost.
This similar pattern has repeated in most technical booms prior to this.
When hard drive technology was improving fast enough that old hard drives were quickly obsolete, IBM could maintain good margins making hard drives. But once hard drives got good enough and advances were slow enough that innovation was not the only factor considered by drive purchasers, commodity hard drives started to take over and IBM had to exit those businesses.
The same is likely to happen once model improvement slows.
by ijidak
7/21/2026 at 9:07:29 PM
You’re absolutely right and it’s heartening to see. I maintain a client with ~every provider you can think of and llama.cpp and it was really tiring the last few days to see people laundering other stuff through Kimi and Qwen. They’re not even open yet, the hype was based on their own blog posts, no one’s actually running these locally, the Qwen Max’s have never been open, Kimi’s API was 1/2 the speed the benchmarks was based on, when it was up, and had 60% downtime before they had to stop accepting new accounts, and their EULAs are “your inputs and outputs are ours.” May being clear-eyed benefit us both in the long run.by refulgentis
7/21/2026 at 9:27:50 PM
"You’re absolutely right and it’s heartening to see"Damnit, I usually don't jump to LLM speech patterns, but this opening had me thinking you were a bot. But after checking your profile, I think you pass as human. I wonder when will be the time, this does not work anymore for me. (Creation date is a strong hint, but abandoned accounts can be hijacked)
by lukan
7/21/2026 at 5:21:01 PM
Yes agreed - I wrote this up a while back https://martinalderson.com/posts/whats-going-on-with-gemini/My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.
They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.
The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.
by martinald
7/21/2026 at 4:35:24 PM
There was some recent reporting that a July release of the Pro model got pushed back for exactly that reason. Its performance was not good compared to the OpenAI/Anthropic big models. They are having a lot of problems with posttrain.by mediaman
7/21/2026 at 4:17:14 PM
This is the feeling i get too. Cant produce quality, but can produce something that is super fast...so take the wins where they are.by verelo
7/21/2026 at 4:51:54 PM
We don't have enough fast models, so I see this as a positive. I just test drove Gemini Flash Lite and it's crazy fast.by copperx
7/21/2026 at 9:19:46 PM
For a coding LLM specifically, when is fast a good tradeoff for quality?by dotancohen
7/21/2026 at 9:25:27 PM
I wouldnt say it is, but there are circumstances when speed is helpful. I wouldn't argue that coding is one of them.by verelo
7/21/2026 at 4:50:03 PM
I personally doubt that.It would be a shame if they cannot beat Kimi K3 or Qwen3.8 Max, both of which are claimed to be Fable-like. If that is true, it will be [or would be] the first time a major American lab falls behind a Chinese competitor.
by maxloh
7/21/2026 at 8:57:57 PM
Google can't compete with China, neither can Meta. Only two labs in the US can keep chucking billions at the frontier race. Everyone else has a real business to run.China can keep up because it's cheaper to run a frontier lab there. They also have more researchers and a stronger cultural inclination for this sort of thing. And I guess the business case in China doesn't have to work as well as it does in the US.
by chrsw
7/21/2026 at 6:30:05 PM
> focusing on what they can get wins in like speed insteadSpeed as a differentiator has always been Google's thing. They (used to?) show the microseconds it took to query & rank web-scale search results. Chrome, notoriously, focused on speed at the expense of resource use. The very many efforts to efficiently speed up Android & its runtime since its inception, and so on...
> their big model underperforms chatgpt 5.6
Possible but TFA claims:
We have started our most ambitious pre-training run yet, for Gemini 4 ...
by ignoramous
7/21/2026 at 9:42:02 PM
That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.by mnicky
7/21/2026 at 8:01:35 PM
Didn't they already acknowledge this?Paywalled article, but the headline is basically all you need: https://www.bloomberg.com/news/articles/2026-07-16/google-ge...
by godwinson__4-8
7/21/2026 at 7:26:17 PM
I choose fourth option.4) googles big model just performs worse than K3 and GLM so they choose not to embarass themself.
Like I love Gemini and use it a lot to one-shot whole MR with huge contexts, but its just much worse when its come to tool use and agentic coding.
by SXX
7/21/2026 at 3:59:05 PM
I wonder if the broad use of AI overviews on Google search results is having an impact. Maybe the numbers make it more profitable to use their compute on several billion searches a day rather than selling API access.by petercooper
7/21/2026 at 4:17:15 PM
I think it's a safe bet that Google seems more interested in making a model that improves Google rather than making a model that improves workers.Fast, light weight, ok intelligence. Perfect for serving 20B+ prompts per day mostly surrounding banal human things.
OAI and Anthropic's cloud spend can cover the revenue gap, as Google is already capturing a large chunk of those guy's revenue.
by WarmWash
7/21/2026 at 5:48:58 PM
Not to mention internal use cases, such as prediction-related tasks like serving ads.by bitshiftfaced
7/21/2026 at 4:21:41 PM
The AI mode on Google search is pretty impressive. Helped me figure out what a bunch of stuff I was seeing out the window was while traveling.by neutronicus
7/21/2026 at 4:58:16 PM
Microsoft also seems to be working in this space. They recently released this:https://huggingface.co/microsoft/bitnet-embedding-0.6b
It’s a small multilingual embedding model designed for things like search, RAG, and semantic similarity. It supports a fairly large context window and is designed to run efficiently on a CPU in a GPU starved world.
The interesting part is that it builds on BitNet, using ternary weights of -1, 0, and 1 instead of the usual floating-point weights. That should make indexing and searching large amounts of text much cheaper without giving up too much accuracy.
by SadErn
7/21/2026 at 5:26:32 PM
AI overview is just a summarization of the top 2-3 results. Of course at Google scale that will still need a ton of compute, but the requirement for generating an overview is many orders of magnitude lower than asking the same question in Gemini.by paxys
7/21/2026 at 8:25:16 PM
based off what?by amazingamazing
7/21/2026 at 8:39:13 PM
vibes (coding)by butlike
7/21/2026 at 9:16:07 PM
More likely they don't manage to advance benchmarks on the SOTA level anymore. In other words: They can't beat 5.6 nor Fableby zwaps
7/21/2026 at 10:05:07 PM
2.5 flash was absurdly capable on a cost basisby joshu
7/21/2026 at 8:53:44 PM
I think it's 2. I frequently get told there's no capacity for Pro and the query is answered by Flash with extended thinking. And tbh it's hard to tell the difference between the two, especially if you're not coding with it.by rjh29
7/21/2026 at 10:12:09 PM
It's hard to tell the difference because they nerfed Pro to oblivion, it used to be much, much better model (even for non-coding/chat)by hn8726
7/21/2026 at 4:54:31 PM
It seems like there are some credible rumors that Google is actually winning in terms of actually building models that work and don't lose money- between how they're able to price them, the TPU advantage and their capex advantage (being able to raise debt + just having a lot of cash - well I said not lose money... more like not go bankrupt).From the outside they look like they're behind in terms of frontier models, but I think they might be the best positioned to not go out of business when the bubble pops.
Also look at the fact that they've been able to deploy AI-assisted search at google scale. It must be another order of magnitude larger (at least) than the model deployments for OpenAI and Anthropic.
Of course unless you're inside Google it's impossible to know for sure.
by awongh
7/21/2026 at 7:02:30 PM
In terms of open models, Gemma 4 beats the pants off everything else to the point that paying for APIs becomes hard to justify. Qwen has the meme-share for coding, but it feels much less well rounded. I have no doubt that Google have both the infrastructure and the expertise to curb stomp everyone else, should they resolve in earnest to do so.Lest we forget, "Attention is All You Need" came from Google.
by dTal
7/21/2026 at 9:18:15 PM
> "Attention is All You Need" came from GoogleIt also came directly from the university of Toronto, and the university of Toronto seeded all American frontier labs (including Grok (why do you think they could start so fast))
by lynguist
7/21/2026 at 9:14:04 PM
Interesting, glad to hear. We have gemma4 at work, and I was considering localhosting qwen, but gemma4 is so far behind the Opus and Fable I have at home that I've decided to hold off for another model release.by scottyah
7/21/2026 at 8:51:01 PM
How long until Gemma 5 hits?by ishurand4
7/21/2026 at 8:55:48 PM
Are you suggesting Gemma beats GLM 5.2?by cherryteastain
7/21/2026 at 10:19:19 PM
At 20x the parameter count I should hope GLM beats Gemma! But is it 20x better? Expertise is demonstrated, not by making big models, but by making small ones. Bigger isn't better if you can't run it at all.by dTal
7/21/2026 at 8:43:56 PM
It's rumored that Gemini 3.5 flash has a >50% margin, and I'd imagine 3.6 flash is even higher.I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting...
I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the race selling $0.50 for a dollar - when everyone else is losing or barely breaking even.
by onlyrealcuzzo
7/21/2026 at 9:15:59 PM
It kind of doesn't make sense though, because typically a large org like Google can afford to crush competitors on pricing. They could probably even go toe to toe with chinese model pricing for years without feeling it.Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?
by WarmWash
7/21/2026 at 7:12:04 PM
That "TPU advantage" might be slowing Google down (though likely not as much as their internal bureaucracy).Porting CUDA-based research, debugging, and overall experimentation speed is likely slower.
The GPU is still king for training.
by deltaqueue
7/21/2026 at 9:09:34 PM
lmao, you know all Anthropic models are trained on TPU right?by anthonypasq
7/21/2026 at 5:58:25 PM
They basically don't exist in the currently most profitable LLM market (coding).Yes, subs like codex are heavily subsidized. But API billing has massive margins and that's what enterprises pay.
by redox99
7/21/2026 at 6:26:52 PM
Does it have "massive" margins? Afaik no one has said publicly what margins there are on an API call?by awongh
7/21/2026 at 8:07:31 PM
"As of October [2025], OpenAI's compute margins reached 70%, up from 52% at the end of 2024 and double the rate in January 2024, [The Information] said, citing a person familiar with the figures."https://www.bloomberg.com/news/articles/2025-12-21/openai-se...
As for Anthropic, the rumors I remember seeing for their API margins were more like 85-90%, but I don't have a reference at hand for those. But once you know the API is wildly profitable and the subscriptions are roughly break-even and not even a big slice of their income, all of the investment makes a lot more sense.
by SyneRyder
7/21/2026 at 5:08:27 PM
Logan Kilpatrick said on an interview not too long ago that flash 3 and 3.5 are the same pre-train. all gains on top of 3 flash are post-trainingby anthonypasq
7/21/2026 at 5:42:40 PM
Maybe, but they said they have “started” the Gemini 4 pretrain. So not having done any significant pretrain in a year or so seems odd to me.by mchusma
7/21/2026 at 9:21:09 PM
Pre-trains take a huge chunk of your compute offline, incurring both an raw expense (24/7 max power for all training clusters) and an opportunity cost (could have sold excess compute during that time). They also don't come with any great guarantees, as lots of techniques look good on small scale and crumble or plateau once scaled.by WarmWash
7/21/2026 at 6:08:59 PM
Maybe it's like Meta not releasing the big version of Llama 4 a year or two agoby ocamoss
7/21/2026 at 4:12:58 PM
I wonder if they waited for the new TPU generation to train a larger base model.by spyckie2
7/21/2026 at 4:40:14 PM
"3.5 pro is testing with partners! will hopefully land soon."by tpm
7/21/2026 at 5:16:43 PM
> the lack of accompanying pro models with these flash releases either means:Rumors say 4) it didn't perform well, especially in coding so has been delayed
by re-thc
7/21/2026 at 4:11:48 PM
Or perhaps 4) it's outcompeted severely by other models & releasing it would only tarnish their nameby jauntywundrkind