alt.hn

7/31/2026 at 1:08:30 PM

Autoregressive Language Model on the 6502 Processor

https://mattbeton.com/blog/bitnet-6502.html

by nmstoker

8/3/2026 at 1:05:16 AM

> The model weights and inference code need to be contained within 25KB of user-space memory

Wouldn’t it be era-appropriate to allow relying on banked memory? You’d still need to hold the inference code, but you could effectively stream(ing page) the weights as you compute on them.

by derefr

8/3/2026 at 2:42:29 AM

From a look at https://en.wikipedia.org/wiki/BBC_Micro#Specifications I think the 6502 versions of the beeb didn't have banked RAM so to keep it loadable from tape the limits might be as stated.

But with substantial additional effort, maybe some banked ROMs could be added..?

by striking

8/3/2026 at 12:59:03 PM

Banking RAM is most simply done by physically wiring RAMs in parallel and routing Chip Select/Chip Enable pins individually from GPIO. People were doing it in PDAs until late 00s. Also, I/O ports on these old machines and CPUs up until switch to PCI Express were philosophically just DRAM slots in alternate shapes, so RAMs could be more or less simply installed and RAMs on adapter cards just accessed. This is also technically applicable to PCIe from various perspectives other than performance.

by numpad0

8/3/2026 at 12:36:34 PM

You could buy a board called Integra-B to expand a model B.

https://www.youtube.com/watch?v=lT07uPRRPfU

A few games required some extra RAM.

Update: Actually, there were other boards. I remember one particular advert for a board that could handle 256KB of RAM/ROM. I'm now actually curious if adding too many ROMs would make the * commands noticeably slower (the OS would look for the commands in the ROMs, one by one, until it found the code to handle them).

by forinti

8/3/2026 at 4:59:00 AM

the banking was done via an external ram controller whose bank control register was mapped to some (unbanked) controller address.

by froh

8/2/2026 at 11:29:24 PM

Cool to think this demo would have been possible over fifty years ago. I wonder what someone from 1975 would have said if you had shown this to them back then.

by bmc7505

8/2/2026 at 11:17:22 PM

This is super cool! As someone who's worked a little with NES programming and tried out cc65, I'm surprised he didn't just hand write some assembly, he likely couldve saved a lot of space if I had to guess.

by tyromaniac

8/3/2026 at 6:17:52 AM

From my experience, modern machine learning models don't scale well down at all, and it's almost certainly better to just use a simple Markov chain variant of some sort - like Niall on the Amiga, or whatever Terry Pratchett used to come up with Foul Ol' Ron's catchphrase, "Millennium hand and shrimp".

by vintermann

8/2/2026 at 11:14:15 PM

The 6502 is notoriously unfit for a C compiler, so probably there is room for more performance in the future. :)

by actionfromafar

8/3/2026 at 1:46:54 AM

The biggest win for AI dev efficiency is cutting down what gets loaded into context. Semantically matching tasks to the top tools helps a lot.

by torment-nexus

8/2/2026 at 11:27:27 PM

This is amazing project! I hope it will result in real miniaturization of AI - for example, edge LLM inside of glasses. That will be awesome.

by toplinesoftsys

8/3/2026 at 4:06:26 AM

really great work

by aghilmort

8/3/2026 at 11:06:18 AM

[flagged]

by jkwang

8/3/2026 at 2:38:30 AM

[flagged]

by leonmeng