8/21/2026 at 5:27:07 PM
For a long time, the chip manufacturers had an inclination to simplify the hardware and rely on the software adapting and optimizing. But for decades this bid failed. Now we have the unrelenting AI capable of finetuning kernels relatively quickly. Maybe simpler hw will work this time? Note: not sure if TPU/NPU is not only simple but also too limited.by empiricus
8/22/2026 at 12:10:25 AM
I have some doubts. For example, there is no software replacement for out-of-order execution: It is unknown before runtime in which cache level the required data will be or which values it will have (which changes the latency of a few instructions).by ahartmetz
8/22/2026 at 12:51:06 AM
What if caching was not transparent, but required explicit management?by Retr0id
8/22/2026 at 1:17:51 AM
That was the PS3 SPU. It was very fast for its time, but only for the small subset of code that could work within its constraints, and viciously difficult to program for.by frogblast
8/22/2026 at 8:19:33 AM
We tried that, doesnt work out in general computers running more than one task.DSPs do that, Atari Jaguar, Sony PS2 and PS3 did, all the GPUs manually manage cache.
More than one task and you start a fight over resources, have to manage hierarchies, priorities, all the stuff that now happens automagically.
by rasz
8/22/2026 at 6:08:32 PM
[dead]by asrgianewrion