@joshfpocock: Someone ran a 2.78 TRILLION parameter model on a single CPU in 8.24 GB of RAM. No GPU, no PyTorch. The whole engine is 176 KB of C. The trick: 93% of the model is experts that never load. They stay on the SSD and only get read when a token routes to them. 5.5 TB at full precision → 8.24 GB resident, byte-identical output. The fine print: 33 seconds per token. Not 33 tokens per second. You also need a 1.56 TB SSD and it’s Linux only. One YouTuber built it and admitted on camera he never generated a single token — not enough disk space. Not a product. A proof. The memory wall was never about model size, it’s about which bytes you keep. #ai #opensource #llm #localai #cprogramming
TPS so low, by the time you get your system prompt loaded, Kimi K4 will be out.
2026-08-18 01:40:19
46
nutrient pets :
.004 tok/sec
2026-08-18 21:34:52
0
Jim :
If you're willing to accept token rates like that obviously you can run a model with not a lot of memory. It's wildly inefficient compared to a system which keeps the tensors resident in memory.
2026-08-18 21:45:23
0
Crazypanda :
Fact check: Bro is lying for short form content. Correction: 93% of the model doesnt have to REMAIN in resident ram, and you can use ur SSD as horrendously slow RAM. The hardware is not what he said either
2026-08-18 15:38:41
13
Jash :
AI data centres are the 1950s computers of the modern age.
2026-08-18 03:48:23
20
nolancongram :
Does this mean ram prices will come down?
2026-08-17 18:01:49
7
Lt Data :
Lol 33s per token is hilarious
2026-08-18 19:36:44
1
🆂🆃.🅱🅰🆁🅽🅰🅱🅰🆂 :
It's fine if you don't want the answer to "What's the weather in my area look like this week" any time this year 🤣
2026-08-18 20:29:16
0
Gabe :
I ran glm 5.2 that way
2026-08-18 02:49:47
3
oXLast2dieXo_TTV :
I run it on 6gb lmao
2026-08-18 01:54:50
2
cat person :
laughable speed but its progress either way
2026-08-18 17:00:44
0
robocop8788 :
❤️kimi and ❤️deepseek
2026-08-18 22:26:49
1
Umang :
how does it know which parts of the model to swap into RAM?
2026-08-18 15:53:30
0
Lumrin :
its probably still generating its first response right now. useless. great concept but let it cook longer.
2026-08-18 06:02:26
8
user888414142131 :
context window: 8 bytes
2026-08-18 11:40:41
6
Humbull :
Why do we data centers then?
2026-08-18 04:23:53
0
redtro.117🇺🇸☢️ :
Ts took 1 year to load
2026-08-18 05:15:13
4
noa_1_24 :
0 to 1 token is 2 weeks!
2026-08-18 01:11:12
3
ky :
no fluff no filler just raw unfiltered
2026-08-18 18:54:09
1
xpppppppppppppppppppppp2 :
bro could someone explain I keep getting these videos lol
2026-08-18 06:09:14
0
Arsenicx2 :
33 seconds a token is why RAM is important. Nothing stops you from running a regular model in SWAP/Virt. AKA on disk it's just going to be dumb slow.
2026-08-18 09:53:28
2
Mike :
I'll keep my 195t/s Qwen Model...
2026-08-18 14:53:48
2
Richard B. :
30s/tok lmao
2026-08-18 06:04:12
2
Cybersecurity GRC :
is this a lie?
2026-08-18 18:44:52
0
To see more videos from user @joshfpocock, please go to the Tikwm
homepage.