How many P40s are you running in parallel? Like 16???
2026-08-18 14:16:26
0
Jared Budlong :
What about a Tesla K80?
2026-08-18 12:55:42
0
user6744183241514 :
It’s been running on a single prompt on my 5080 64g dram for like 40 hours!! It’s crazy
2026-08-18 13:49:12
0
Victor :
I can run this on my 3060 12gb. it's the first time local models are actually useful and not just a fun toy. The Q2xss runs are usable speed, it's slow but no token limits so you if you don't have a ton of money to burn on paid models like me it's the only thing you can run for hours non stop
2026-08-18 14:11:18
0
Tom Wilson :
It might be the first time where I felt like I could give up my Claude sub
2026-08-18 12:32:34
2
LeftyVeteran :
why do I feel like this is the repeat of when IBM charged memberships to mainframes and then the PC dropped and mainframe access became worthless
2026-08-18 04:50:47
45
thtnigeriankid :
I can only imagine what Qwen 4.0 27b will perform like. Local AI is truly the future.
2026-08-18 09:47:21
7
ComradeWilly :
I just installed the p40 you told me to buy last year so this release is fortuitous
2026-08-18 11:31:59
2
David :
I have a strix and I’m only getting 55k context window.
2026-08-18 05:38:29
1
Brian 🇺🇸 :
Nice, I'll try on my P40s.
2026-08-18 13:15:03
0
BrockMcBreadcat :
ive been using it and its been really good
2026-08-18 12:02:36
0
Matt D :
America has been dumping money into AI to fund oligarchs, China has been funding AI to win.
2026-08-18 06:11:23
12
Zuck :
2T/s but for non coding tasks is good
2026-08-18 09:47:32
0
Roald over :
running it on an arc b70 pro 32gb vram, q6_k_xl 128k context. it really is very good.
2026-08-18 05:45:09
3
user2566897059072 :
On a p40 how much context would you be able to fit?
2026-08-18 05:35:53
0
Matthew Stone :
I have 2 x 3090s I bought a long time ago and 96gb ram and I’m excited. Planning to use my ChatGPT data export to populate an obsidian for memory.
2026-08-18 08:50:48
3
FrufruKachuu :
Its very good! But there is a performance penalty compared with 3.6. In my agent its about 2-3x slower on the same reasoning effort. It seems to yap a lot more 😅
2026-08-18 08:45:15
0
lol k :
Thought I needed that much in vram for these models, but I’m sitting on 128gigs of ram before everything jumped in price. Think I might just load this up finally.
2026-08-18 04:47:28
2
Mel :
Deepseek has better latency because of the caching
2026-08-18 05:17:07
0
Stranger :
I used it yesterday for a specific and kind of rigorous inference series I'd been using opus 5 for up to now and for the first time local model response was superior. it seems really excellent.
2026-08-18 06:53:49
2
ToldYouBro666 :
It's a very good model, but for example I have a 3080 with 10 GB of VRAM and 32 gigs of ram and I can run it, but it's really not usable. Too slow and makes the computer completely unusable in the meantime. It's getting to the point where it might be worth it to buy separate pc for running stuff.
2026-08-18 06:22:31
0
BAF1's :
Is it good enough to break out of a sandbox?
2026-08-18 05:00:12
2
Q :
petite question: wont it be very slow if you try to run it off of an SSD with internal computer ram less than 24Gbs of RAM? if I had a laptop 16gb of RAM but 4tb of SSD, do you think it would still run efficiently?
2026-08-18 06:03:32
0
2scoopsplz :
I have a 5080 with 16gbs of vram and 64gb of ram. Can I run it?
2026-08-18 07:39:50
0
CD :
But ram is going to keep getting more expensive 😢 I should have bought more ram when I built my computer a few years ago.
2026-08-18 04:19:27
1
To see more videos from user @cjtrowbridge, please go to the Tikwm
homepage.