Claude told me I would get only 20 tokens a second with a Qwen 2B model on a OnePlus 7t with 8g ram. why is it wrong
2026-09-11 18:44:00
0
Heathrow | Not the Airport 🛩️ :
How I feel running local models 24/7 and keeping my money from OpenAI and Anthropic.
2026-09-11 17:54:03
26
BigScentEnergy :
Could you run this on a hetzner server ?
2026-09-11 23:03:34
0
Stove :
how do you like it in comparison to DeepSeek v4 pro?
2026-09-11 22:31:23
0
virtualuman_ :
I Dont Belive You! lol
2026-09-12 00:37:43
0
Josh Wren :
I ran it on a computer with 4gb vram gpunand got 2 tokens per second
2026-09-11 23:58:04
0
Mechanical Tangerine :
Qwen 3.8 is supposed to be able to run on my machine, but I’ve had zero luck so far. I have an M5 with 24GB RAM, which isn’t great, but it’s way better than half the machines you listed in this video
2026-09-11 20:41:47
1
joshandfound :
Resilience has value for sure, complex tooling paired with batching makes sense!
2026-09-11 23:26:33
0
charliecearbhaill :
I am struggling to find a worthy thing to do with my local ai, I use free cloud chats for my day to day and it does what I need for that so using it while it's free. but need to figure out something to give to a model to chug on locally. I think I lack the creativity
2026-09-11 17:51:10
3
Close To 7734 :
what's the trade off
2026-09-11 22:24:51
0
Victor Megir :
but does it work on an orange pi 6?
2026-09-11 21:54:06
0
bw201417 :
do you have any recommendations for a guide or video resource for initially setting up agents locally? I have the worst perfectionism anxiety about needing to research all the things to make sure Im doing it optimally and don't F up my computer. esp with the incidents of AI breaking out of their containers. Idk if the big companies did something reckless to make that happen 🤷♀️I feel like I need a human resource than just asking AI how to do it so I know my initial set up is solid.
2026-09-11 20:08:06
2
J.C. D. :
When Deep Seek jacked up their rate I tried switching my workers to Ali Baba /Qwen but it kept disobeying my guard rails. Wouldn’t recommend for builder planner or validator
2026-09-11 20:03:10
0
Trent Herring Khaos :
I’ve been running it off a Rtx 2070 115watt, it’s only like 5 tokens a second but this model is awesome! Zero refusals using the obliterated Hua Hua model 🤘🏿
2026-09-11 22:34:06
1
Great Trousers Mate :
CJ what's your view on the increasingly apocalyptic language regarding LLMs from industry insiders? I find myself feeling sceptical, but I also don't know enough about the subject to have a confident opinion.
2026-09-11 17:58:20
2
bo :
can it be easily finetuned?
2026-09-11 18:52:08
1
halkpaw :
Does it have tooling similar to Claude code or codex? Last time I tried something like this, it kept messing up with tooling and wasn’t worth
2026-09-11 21:04:59
0
sin :
Do you think there is a real possibility they ban open source models?
2026-09-11 17:58:58
1
magicada :
you can't get the cerebras speeds. its input token limited (so turn limited) by a lot.
2026-09-11 20:27:17
0
alexanderschmitt9 :
What do you think about running these models on M4 Max MacBook pros? What model would you run on a 32gig variant?
2026-09-11 21:00:52
0
communist bobby bacala :
hmmmmmm instead we should spend a trillion dollars on building data centers to run models that basically do the same thing
2026-09-11 18:41:26
0
Eric Orr :
I've been building a sort of cobbled together homelab with my old laptops running qwen coder to build my first GitHub webhooks. So far I've been pleasantly surprised, and it's fine that it's slow cause it's just supplementary workflows.
2026-09-11 21:55:22
1
DigitalNeal :
Kicking myself for buying a 12gb 3070 7 years ago 😅
2026-09-11 19:19:54
1
Brian 🇺🇸 :
I love my P40s. They are workhorses.
2026-09-11 18:35:48
2
Delilahforrealio :
It took 20 minutes for it to respond to “Thank you” 🙄
2026-09-11 19:05:57
1
To see more videos from user @cjtrowbridge, please go to the Tikwm
homepage.