8/1/2026 at 11:22:08 AM
Idk about you guys but I'd find 0.5t/s useless. Even for long tasks.I'd rather just shell out the money to offload as much as possible to say 2x 4060ti 16gb with tensor parallelisation. Anything but that low token rate.
This is the sort of thing I'd expect in 20 years for some cyberpunk esque "turtlebot" that thinks at 0.5t/s, is solar powered and performs some menial civic maintenance background task like cutting grass, or scrubbing pavements. Or the "slowbot" that sits in the garden slowly pruning a bonsai, only just keeping up with the growth of the young plant.
by fennecbutt
8/1/2026 at 2:06:35 PM
Yes it’s mostly unusable but ongoing iterations of projects like this will eventually lead to a usuable version so good to seeby fsuts
8/1/2026 at 4:01:02 PM
I pay $20 for codex, use it daily for coding, and still haven't dipped below 50% for weekly usage.I wouldn't even be able to afford buying 29 gigs of ram, or a new video card with hardware prices the way they are now.
Maybe if it was 2016-2018 prices, I'd think about it.
by preommr
8/1/2026 at 8:28:39 PM
Well, if you only ever needed a $20 subscription's worth of tokens it would never make any sense to buy hardware since even at 2018 prices you could buy 20 years of subscriptions.by root_axis
8/1/2026 at 9:34:43 PM
Assuming $20 subscriptions continue to be available for 20 years. Maybe they will, I don't have a time machine, but $20 subscriptions seem to be the $1 Uber ride level of pricing.by fragmede
8/2/2026 at 1:04:59 AM
Sure, I'm just saying, even for "mundane" computing tasks where the hardware is cheap, if you're optimizing for cost, renting hardware is always the better choice.by root_axis
8/3/2026 at 2:00:01 AM
It could be handy if it's some batch work you don't need quickly. But that's pretty rare, tbh. It's way too slow for interactive stuff. With this kind of model I like to spar, fire ideas at and get feedback. Having hours of delay every time doesn't work for my ADHD brain.by wolvoleo
8/1/2026 at 7:50:25 PM
It's about a million tokens a week. So you can do the usual calculations of rent Vs buy.There are definitely some tasks, that if the system can run unsupervised (a largeish if), it doesn't matter as long as the result happens before a deadline.
In that respect it is easy to tell if this works for you or not. As a stepping stone to more efficiency in the future it has more value.
by Lerc
8/1/2026 at 8:23:25 PM
That’s a good way to think about it - do you want your expensive laptop grinding away constantly 24/7 consuming $5 of power to generate what amounts to $15 worth of tokens per week ?You can get a frontier model subscription for about the same cost as the electricity (heavily subsidised by someone else’s money !) with instant results.
I know there are applications for this and it’s cool people are pushing the boundaries with local models. I run oMLX and Qwen myself, but it’s a toy really. It’s not anywhere near replacing Claude for my purposes.
by danw1979
8/2/2026 at 12:32:19 AM
Would it really be $5 of power? I know it's hard to say definitely but even over a week that seems high to me for Apple Silicon. Don't those processors sip power?Probably a useless comparison, but in my NYC studio apartment, my electric bill is usually ~$75 per week during the summer. That includes air conditioning, appliances, etc etc, not to mention the Haswell Desktop PC I keep running 24/7, which is rarely idle because I queue up compile jobs and stuff.
by Wowfunhappy
8/1/2026 at 5:48:16 PM
I mean, yes, it’s for sure too slow. If you want something faster you can run Qwen models locally at a very decent speed. Running Kimi K3 with 29GB of memory is pretty impressive in itself, but is more an experiment than something supposed to be viableby dgellow
8/1/2026 at 7:29:46 PM
paying for the hardware is also too expensive for no reason, you're not querying the llm 24/7. Better use a service that hosts a lot of popular open weights LLMs, unless you have a very good reason not to.by make3