7/22/2026 at 11:03:42 AM
Doesn't really explain why its faster. "Optimizing for every kind of CPU" is not really enough info.Cool project nonetheless, I will go through the code later tomorrow
by ramon156
7/22/2026 at 12:13:21 AM
by convexstrictly
7/22/2026 at 11:03:42 AM
Doesn't really explain why its faster. "Optimizing for every kind of CPU" is not really enough info.Cool project nonetheless, I will go through the code later tomorrow
by ramon156
7/22/2026 at 12:14:09 AM
GitHub: https://github.com/marcelroed/gigatokenby convexstrictly
7/22/2026 at 9:58:13 AM
They even added a section ”AI Use Disclosure”, beautiful!by Alifatisk
7/22/2026 at 10:06:48 AM
I kind of assumed the model would process the text 'directly', from what I understand, wouldn't this be biasing the input based on how you tokenise as it's lossy?I assume this tradeoff is purely for speed/compression. Or am I missing what's going on here?
by benj111
7/22/2026 at 10:33:59 AM
Tokenization is done on the CPU. Models never see the raw characters. That's why you get trick questions like the number of r's in strawberry.There are many research papers on models using characters directly. One challenge is that effective context length is smaller.
by convexstrictly
7/22/2026 at 7:31:56 AM
[dead]by doosdom