@nawraskader: A 2.78-trillion-parameter model just ran on one CPU with 8.24 GB of RAM. The catch: Kimi K3 did not shrink to 8 GB. Its 1.56 TB checkpoint stays on disk. Because Kimi K3 is a mixture-of-experts model, only 16 of 896 experts activate per token. A 176 KB C99 engine streams those expert weights from NVMe and moves the 108.81 GB trunk through RAM one layer at a time. No GPU. No CUDA. No framework. But this is a capacity proof, not a practical chatbot setup: the 8 GB run takes about half a minute per token and still needs Linux x86-64, AVX2 + FMA, and about 1.7 TB of fast local storage. The engine is Apache-2.0 open source. Kimi K3's weights remain under Moonshot's separate Kimi K3 License. Save this if you care about local AI systems—and subscribe to Bleeding Edge AI for AI, software, and entrepreneurship without the noise. Link in bio. #AI #OpenSource #LocalAI #MachineLearning #KimiK3
If the proof of concept works, this will be the slowest it will be. Few weeks. Few months. And we will have a useable framework. Open source FTW
2026-08-10 01:17:39
1
Stanton :
I’m genuinely curious to see a large quant of of this, like maybe half the size, using the same tech. I have a 3090 w 24GB VRAM. So, 4x what is required for this, with half the size maybe… maybe… could be getting close to usable? I mean, that might be close to Sonnet, which is usable for a LOT of work. I’d crawl across broken glass for that, TBH.
2026-08-10 02:01:31
1
To see more videos from user @nawraskader, please go to the Tikwm
homepage.