alt.hn

8/21/2026 at 4:16:42 PM

What happens when a GPU reads memory

https://blog.doubleword.ai/what-happens-when-a-gpu-reads-memory

by ibobev

8/21/2026 at 5:27:07 PM

For a long time, the chip manufacturers had an inclination to simplify the hardware and rely on the software adapting and optimizing. But for decades this bid failed. Now we have the unrelenting AI capable of finetuning kernels relatively quickly. Maybe simpler hw will work this time? Note: not sure if TPU/NPU is not only simple but also too limited.

by empiricus

8/22/2026 at 12:10:25 AM

I have some doubts. For example, there is no software replacement for out-of-order execution: It is unknown before runtime in which cache level the required data will be or which values it will have (which changes the latency of a few instructions).

by ahartmetz

8/22/2026 at 12:51:06 AM

What if caching was not transparent, but required explicit management?

by Retr0id

8/22/2026 at 1:17:51 AM

That was the PS3 SPU. It was very fast for its time, but only for the small subset of code that could work within its constraints, and viciously difficult to program for.

by frogblast

8/22/2026 at 8:19:33 AM

We tried that, doesnt work out in general computers running more than one task.

DSPs do that, Atari Jaguar, Sony PS2 and PS3 did, all the GPUs manually manage cache.

More than one task and you start a fight over resources, have to manage hierarchies, priorities, all the stuff that now happens automagically.

by rasz

8/22/2026 at 6:08:32 PM

[dead]

by asrgianewrion

8/21/2026 at 6:39:19 PM

Can someone help me understand why I spent 5 minutes reading something that I still don't understand?

Link for the ELI5 version?

by snigacookie

8/21/2026 at 9:05:48 PM

There isn't really an ELI5 version of computer architecture but you could start with the book Inside the Machine by Jon Stokes. Then you can get into SIMT.

by wmf

8/21/2026 at 8:49:17 PM

> Little of the detail of this path is documented by NVIDIA, at least not to the level that we’d like, so we’ll determine it by running timing experiments on the hardware itself

Or you could just use the AMD isa.

by xyzsparetimexyz

8/21/2026 at 6:00:09 PM

This is a type of article that is HN worthy and came to HN for initially because I don't even understand a third of content there. Giving me inspiration to dive deeper.

SEriously, I don't understand it (yet) lol.

by yipinwong

8/21/2026 at 7:02:23 PM

That's the best feeling -- idly clicking through looking for that one rabbit hole to fall down then stumbling upon a gem like this. I noticed this the first time on copetti.org articles about game console architectures: know nothing, look every jargon or acronym up as you read along, end up with 42 tabs and a basic high-level understanding of the topic by the end.

by nazgulsenpai

8/21/2026 at 6:20:02 PM

Ctrl+F PCIe BAR.. nothing.

by brcmthrowaway

8/21/2026 at 6:24:00 PM

This is talking about HBM/GDDR

by porridgeraisin

8/21/2026 at 6:30:36 PM

Something needs to be transferred from sysmem right?

by brcmthrowaway

8/21/2026 at 7:19:28 PM

That's probably PCIe DMA initiated by the GPU so I don't think BAR is used.

by wmf

8/21/2026 at 8:58:01 PM

Pcie dma uses BAR registers setup by cpu. It will contain information about the physical memory address for one. On newer systems, there is NVLINK C2C which is more tightly integrated and less general.

Regardless, my point was the the article is about vram.

However, there is one situation when vram access itself uses the bar, to be fair. When you do P2P dma, code (kernel, either the inbuilt version or the one in the nvidia driver) running on the cpu sets up the DMA engines's GART to contain the BAR1s of the other.

by porridgeraisin

8/21/2026 at 6:28:00 PM

a.k.a. VRAM

by KK7NIL

8/21/2026 at 6:54:00 PM

YMMV; not all GPUs work exactly like this

by mathisfun123