@ai.honeycove: LLM just made running 70 billion parameter AI models possible without supercomputers by loading layers from your hard drive. Its flash attention keeps memory low, letting complex models like LLAMA 3.3 70B run smoothly on MacBooks and gaming PCs. A game-changer for students and indie devs exploring AI. #llm #llama33b #aiinnovation #opensourceai #machinelearning
do I still pay for token usage on a open source llm?
2026-09-17 05:16:04
0
Renzo :
What would this do to a hard drive though
2026-09-17 05:21:03
0
statisticalvoid :
10 seconds per token
2026-09-17 05:24:11
0
super nabilion :
bla bla bla you need huge gpu 24g ... impossible
2026-09-17 05:10:18
0
DogsAreTheBest :
one token per sometimes?
2026-09-17 02:46:31
12
petar⚡️ :
one token per some years
2026-09-16 23:03:12
13
Bon Beau Boudreaux :
everyone complains about token speed. maybe you need to ask it the most precise and comprehensive question using a very precise format with output in a particular compressed format. think about it. what if your output was version that was able to compress several responses into a few very small token group?
2026-09-17 03:48:55
1
androiduser525 :
I already run this llama 70B locally on my asus gaming laptop with 64GB RAM and 8 GB vram. it runs slow. I dont know that this method would run faster
2026-09-17 02:52:40
0
Ryan :
/yawn, so last month, collibri is far better,
we have massive ai compute and we're building shit in python? sigh
2026-09-17 02:10:57
2
Paul Grinstead :
One token per century.
2026-09-17 03:17:34
0
DimiDi :
Ai companies going broke +Nvidia
2026-09-17 03:47:05
0
KBN :
2026-09-17 00:43:18
0
To see more videos from user @ai.honeycove, please go to the Tikwm
homepage.