7/24/2026 at 3:10:47 PM
You can't parse [X]HTML with regex. Because HTML can't be parsed by regex. Regex is not a tool that can be used to correctly parse HTML.by danlitt
7/24/2026 at 4:29:51 PM
Some commenters are missing that this is a reference to https://stackoverflow.com/a/1732454.by layer8
7/24/2026 at 4:57:13 PM
> Locked. There are disputes about this answer’s content being resolved at this time [sic]And also
> Nov, 2020
And then, StackOverflow asks itself why it looses users.
by chrisandchris
7/24/2026 at 4:59:18 PM
The answer is actually from 2009. As long as HTML/XML doesn’t suddenly become a regular language, I think it’s pretty timeless.by layer8
7/24/2026 at 5:07:38 PM
It's a 17 year old joke, maybe one of the most famous on stackoverflow for the aging engineers out there. No one wants new "funny" edits on it.by magicalist
7/24/2026 at 4:35:20 PM
Good ol' appeal to emotion devoid of technical argument.by pwdisswordfishq
7/24/2026 at 3:21:05 PM
You can parse a subset of it though, like if you're in control of the html yourself and avoid certain structuresby rokkamokka
7/24/2026 at 3:32:17 PM
Parsing HTML with a regex is never a good option, but it's sometimes the only option.by zarzavat
7/24/2026 at 3:43:01 PM
In the example from the article it certainly is an option. In Python you could either use a "soup" library or you could play around with a tool like https://www.w3.org/Tools/HTML-XML-utils/man1/hxpipe.html.The more fundamental question for me is why the author didn't decide to make make code blocks non-breaking by default, or just add the class annotations when he writes the HTML?
by pkal
7/24/2026 at 4:24:48 PM
You can absolutely parse HTML with regex, so long as the document is finite in length. Every finite language is regular, hence can be parsed with regex's.by tyho
7/24/2026 at 5:05:53 PM
Zalgo and situational subset parsing aside:> You can absolutely parse HTML with regex, so long as the document is finite in length
This isn't sufficient, unless I'm misinterpreting what you're saying. It's not enough to have documents of finite length (all documents are finite in length), you need documents with a max length, so you have a finite number of possible documents to parse.
by magicalist
7/24/2026 at 4:30:13 PM
Do a web search for the parent of your comment, read the Stackoverflow answer. It's a classic. Learn about Zalgo and Tony the pony, he comes.by gpvos
7/25/2026 at 4:36:37 AM
and you write your regex specifically for that document length, handling all possible nesting combinations. combinatoric explosionby vrighter
7/24/2026 at 3:40:58 PM
Tony the pony, he comes.by matheusmoreira
7/24/2026 at 4:19:45 PM
What do you mean you can't. I do it all the timeby timedude
7/24/2026 at 4:36:38 PM
You do some parts all the time.by throw1234567891
7/24/2026 at 4:34:56 PM
I do not count substitution as "parsing"by prmoustache