@github.awesome: Fitting a 27-billion-parameter model onto one 24GB gaming card usually means stopping at "it loads." This recipe pushes past that to 417 tokens a second batched, and the wins are unglamorous in the best way. Qwen's untied embeddings ship as two unquantized 2.5GB matrices nobody bothered with, so requantizing both to int8 hands back 2.6GB, and on this architecture spare memory converts straight into batch size. #github #opensource

GitHub Awesome
GitHub Awesome
Open In TikTok:
Region: US
Tuesday 01 September 2026 01:41:39 GMT
18185
525
18
90

Music

Download

Comments

statisticalvoid
statisticalvoid :
417 TPS? I don’t believe you
2026-09-12 10:33:33
1
whateverconsciousnessis
whatever consciousness is :
417 is crazy
2026-09-02 16:59:13
0
dafkaoo7
Dafka Gaming :
I had nothing but trouble with this model. Bad code and endless loops and thinking loops no matter what i did 😔
2026-09-06 19:13:54
0
fpvdad
FPVDad :
I’ve got a 3090 24gb with 128gb ram spare and added 3.8 27B via LM Studio, where can I find out more about this? I’d really like to get my tokens up.
2026-09-23 22:46:45
0
bendrinking
bendrinking :
I got it running on a nanopi with 32 gb of ram utilizing the NPU w8a8 dense 27B. Blazing 1.09 toke s per second🤣 It answer very well
2026-09-05 08:51:16
0
hondogamer
HONDO :
Lobbying has started to outlaw this 🤣🤣🤣
2026-09-03 04:06:12
3
user000203144
Y A :
You can fit QWEN3.8 27B on a 5080 laptop (16GB VRAM) without spilling into RAM with over 200k q4 KV context. Check out GSQ-RCO. Can even run DFLASH2 if I reduce context a bit.
2026-09-20 22:07:52
0
rodja1996
Reinhold „ErrOrGamep :
i am running Qwen 3.8 in 16gb vram.
2026-09-19 12:12:54
0
bigfriendlymaker
aaronanderson23510 :
Hmu when it's 16gb vram
2026-09-01 01:57:03
4
imthatmike
Mike :
😳😳😳
2026-09-18 12:17:27
0
To see more videos from user @github.awesome, please go to the Tikwm homepage.

Other Videos


About