7/26/2026 at 12:45:52 AM
The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).
by simonw
7/26/2026 at 3:49:03 AM
Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly.That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
by dannyw
7/26/2026 at 9:43:50 AM
The irony here is that the people blocking all other crawlers are the ones shoring up their monopoly. If you can't block Googlebot because you need the search traffic but you block everybody else so that nobody other than Google can index your site, how do you expect to ever get any search traffic that isn't from Google?by AnthonyMouse
7/26/2026 at 2:44:17 PM
Search traffic is nose diving due to LLM use. The fundamental calculus with google is you install GA and it helps your SEO has changed and google is riding out what will eventually wither as people start to reevaluate the trade off.by tempest_
7/26/2026 at 7:47:36 PM
It's nose diving due to Google putting the LLM answer box at the top of the search results. But then isn't it even worse to be blocking every search engine crawler that isn't doing that?by AnthonyMouse
7/26/2026 at 5:04:06 AM
I think it's bad, because everybody is desperate to hold onto every last bit of google search traffic they can, so they're going to allow training to do so. Google's predatory, unfair and illegal actions will continue as they have with a few $100 million slaps in the wrist from the EU and a few more white house dinners for their CEO.by troyvit
7/26/2026 at 1:40:26 PM
Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index.> Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. https://developers.google.com/crawling/docs/crawlers-fetcher...
Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.
by Deathmax
7/26/2026 at 1:48:39 PM
AI crawlers, famous for respecting robots.txt ;)by mysterydip
7/26/2026 at 11:46:32 AM
EU will write some strongly worded letter.Saying this as a European who is pro EU.
Why should they?
by Scroll_Swe
7/26/2026 at 6:38:14 AM
I don't think it's the Google bot DDOSing people's infrastructure for AI training...by PunchyHamster
7/26/2026 at 8:32:13 AM
[dead]by xhi
7/26/2026 at 12:13:56 PM
This had me in disbelief since the minute I saw it: Google's "AI overview" presumably trained on content from other websites, disincentivizes users from clicking through to those websites..How is that not conflict of interest??
by Razengan
7/26/2026 at 3:11:12 PM
> mandating owners give consent for AI training or drop out of GoogleHuh?
Search and AI are hand-in-hand.
They both rely on embeddings. (Unless you still do keyword-only search, but that's not as good.)
by paulddraper
7/26/2026 at 11:04:04 AM
How come you say "Good." and you are near the top of the comments but I say "Good." and get flagdead?by inigyou
7/26/2026 at 11:39:28 AM
it’s likely to do with the fact that the parent comment here laid out a thoughtful basis / argument for their “good”, providing some detail to their justification for it.that’s just my take/feedback, take it or leave it. i won’t be engaging further as i already feel i’m going against the site guidelines with this!
by dijksterhuis
7/26/2026 at 1:01:35 AM
We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.by jofzar
7/26/2026 at 5:27:50 AM
The vast majority of Googlebot user agents are lying. Real Googlebot is pretty well behaved in my experience. You should use reverse DNS or IP lists to check: https://developers.google.com/crawling/docs/crawlers-fetcher...by bobbiechen
7/26/2026 at 1:53:39 PM
If you're using Cloudflare, set up a security rule to block requests that have "Googlebot" in the UA and are not recognised by CF as a real bot.by simondotau
7/26/2026 at 2:04:14 AM
Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delistedThis is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this
by 20k
7/26/2026 at 5:21:35 AM
If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservativeAlso I should note there are lots of fake Googlebots...
by weird-eye-issue
7/26/2026 at 6:23:34 AM
Indeed, I've got a site which gets a lot of bot traffic and google bot is pretty sensible compared to a lot of other mainstream bots.by remus
7/26/2026 at 9:24:30 AM
Yes, it really does not make that many requests. In fact lots of site owners struggle with having it not crawl and index their site enoughby weird-eye-issue
7/26/2026 at 11:59:17 AM
It is mostly, but it doesn't take a lot of googling to find sites getting ridiculous amounts of traffic from googlebot on google IPs. Its one of the most common complaints about google's search indexingby 20k
7/26/2026 at 1:06:18 PM
Lots of people abuse Google Cloud to get a "Google IP" for a fake Googlebot. Why don't you show me a single screenshot from Google Search Console showing a high number of requests to a site that would be counted as a DoS? All requests from the official Googlebot are logged there so if it's such a common problem it must be very easy for you to show me this.by weird-eye-issue
7/26/2026 at 9:50:28 AM
Don't that feel like a threat to businesses who dare to avoid their content being stolen?by motbus3
7/26/2026 at 5:30:14 AM
People will use something for search and something needs to index pages, either for LLM or old school search engine.by miohtama
7/26/2026 at 1:17:39 AM
[flagged]by inigyou
7/26/2026 at 1:51:59 AM
Why should I use something other than Cloudflare pages for a simple app landing page?by Cider9986
7/26/2026 at 9:16:30 AM
Because you value the internet being decentralised.by inigyou
7/26/2026 at 3:31:25 AM
Because your viewer/customer base will be reduced.by ipaddr
7/26/2026 at 6:52:36 AM
It would have a domain. It affects it even then?by Cider9986
7/26/2026 at 9:16:51 AM
Well, now Google won't be able to see it.by inigyou
7/26/2026 at 1:54:10 AM
Why is this?by ajmurmann
7/26/2026 at 1:59:45 AM
It's a planet-scale MITM?by ceejayoz
7/26/2026 at 2:38:12 AM
It's a cache. My tiny websites couldn't survive getting hammered by AI bots without them.by TurdF3rguson
7/26/2026 at 9:11:16 AM
Are you sure? Have you tried, or did Cloudflare just tell you that?by inigyou
7/26/2026 at 8:51:44 PM
I wouldn't need a cache if my $6 server could handle 1M hits a day.by TurdF3rguson
7/26/2026 at 2:30:55 AM
So you're against all CDNs?by dbbk
7/26/2026 at 2:39:41 AM
A CDN doesn't necessarily have to perform a MitM. We really need more nuanced terminology to distinguish the various approaches.by fc417fc802
7/26/2026 at 2:53:46 AM
Right, but practically speaking all CDNs are MITMs. If you're against cloudflare you should be against cloudfront, akamai, etc. as well.by gruez
7/26/2026 at 9:12:38 AM
Cloudflare is egregiously bad because of its marketing strategy. It tried to get everyone with any small website to use it, by selling a vague notion of security and charging no monetary price, and it worked. They'll even sell you a domain name to increase lockin. Many people recommend getting domains from cloudflare because apparently they're cheap.Akamai, Fastly, etc only take big customers who know what they're doing. You need to sign a proper contract with them. They aren't low-friction.
by inigyou
7/26/2026 at 2:59:46 AM
Ideally yes, the TLS termination does not need to happen for caching purposes. Challenge is that in practice every business wants to be sticky and try to provide more functionalities which do require TLS termination. Most people either trust CDN's or they do not understand MitM so it does not concerns them. Plus they are getting certificate management and DDOS prevention capabilities.by sandeepkd
7/26/2026 at 2:54:00 AM
How would they cache and serve responses without decrypting the traffic?by edaemon