@momus.ai: RTX 5090 QWEN 3.6 35b 800 tok/second via ninfer. Insanely fast #qwen #localai #aiforbusiness #nvidia #5090

Momus - Local AI Lab
Momus - Local AI Lab
Open In TikTok:
Region: US
Friday 11 September 2026 05:17:40 GMT
17082
407
56
32

Music

Download

Comments

quandinglus
Quandinglus Lemar Tickleton :
unfortunately 3.6 is not that smart. Wish you could get this kind of speed for 3.8 27b
2026-09-11 18:32:18
19
nidofdanes
NidofDanes :
I would expect that sort of performance for $5k
2026-09-11 11:51:30
25
thgandalf
thgandalph :
yeah that did not happen 😁
2026-09-12 06:41:54
0
zabwie
Zabi 🦇 :
I can’t wait for them to drop a local MoE qwen 4. I’m running that same model on an rtx 3060 and 32GB RAM. But hear me out (no it’s not vibecoded, yes I have a repo), I’m running at 20t/s at ONE MILLION context. I’m using freetoken + turboquant + my own research + other mini research paper’s online (expert cut and 2 other which I don’t remember), I also rewrote hoe to model saves kv cache to make it faster and use less computation. Recall is perfect at 1M but when I do optimizations like vectorization, it fails the needle test 😭 been at it for over a day
2026-09-11 16:34:22
4
damondanieli_
Damon Danieli :
Nice! That’s nearly 10x what I got for that model on my DGX spark. What are you getting for the 3.8 27B? (I’m running 3.8 flash next on two sparks and loving it — super thorough on xhigh)
2026-09-11 08:44:07
8
rolfbringerofdarkness
Rolf :
trying it out today on my 3090 with Qwen 3.8 27b
2026-09-11 14:37:08
0
jasontorres5603
MoonUnit :
This is madness to pay this. I'd rather spend the 200 a month to get near unlimited usage and try to make money than spend 10k on a few cards that won't be worth it in 3 years.
2026-09-12 02:23:28
0
macroni_lime
Maximus :
tok/sec is cool but I'd be curious to see what kind of value it can generate compared to a frontier model. what it can and can't do
2026-09-11 19:05:53
3
gudangbaju30
Gudang Baju :
that is so fast
2026-09-11 14:55:14
1
t5776811
Mike :
its usually quick when you only process 3B weights, the 35B qwen is not a dense model
2026-09-11 11:02:01
1
weather.tk
Weather Tk :
me with my 3060 trying to run qwen 3.8 27b with token per sometimes in q2
2026-09-11 20:18:24
3
itzpezz
itzpezz :
Huh? What? What are we talking about?
2026-09-11 11:41:42
1
chickenchaser97
ChickenChaser :
What is the stack to achieve this speed and what context are you running it at? I was capping out closer to 200 t/s but im sure my stuff isn’t configured or optimized correctly
2026-09-11 15:16:45
2
paladinpoet
paladinpoet :
It doesn't say how not useful the performance is unfortunately [Teary eyed] but sick speed
2026-09-11 19:23:11
1
culinary.time.machine
Daniel :
2026-09-11 15:06:51
2
lvmrevolutiondotcom
lvmrevolutiondotcom :
si ok fevi far vedere la generazione della risposta...😂
2026-09-11 13:41:42
1
sharpclawser
sharpclawser :
how bruh?
2026-09-11 10:34:13
0
onecheesesteakplease
CheeseSteak :
still sucks
2026-09-12 00:48:13
0
casey00087
Casey ❌ :
😂😂😂
2026-09-12 03:16:28
1
To see more videos from user @momus.ai, please go to the Tikwm homepage.

Other Videos


About