@marcinteodoru: Full setup is in my profile. This is Colibri, and it can run absolutely massive AI models locally by using your SSD, RAM and VRAM together. We’re talking 744 BILLION to 2.8 TRILLION parameters on consumer hardware. It’s still slow right now. But if this gets faster, local AI gets very interesting very quickly. And NVIDIA might want to pay attention. 👀
Not a loop hole… more like a compromise. It’s actually the methodology computing has ran on a very long time, which is very inefficient for the purpose… hence new data centers
2026-09-15 01:09:12
0
stacy lindsey :
What model would be best to run on a ram heavy server build.
2026-09-15 00:52:57
0
Ten Ten :
ok but howany tokens per sec?
2026-09-14 20:31:30
1
Nathy Mosquera :
You do understand that it remains a hardware limitation issue? This is putting a moped engine in a tow truck. And then hopes the moped engine will be more powerfull. Even normal ram is not efficient (slow). Cpu power is also slow. Gpu hardware is very efficient but also expensive. So even Vram is not the issue, even a rtx 2060 with 1 terabyte vram will be slow for this model, even a 5090. I hope you understand.
2026-09-14 20:57:12
5
MarshMellow :
I always thought it was pronounced Mar·tchin
2026-09-15 02:44:17
0
⛔️NoContentOnlySignalComment :
I am tired of this kind of advertisment of some stupid product that just use API of gpt and sell you for 20-50 USD
2026-09-15 00:53:31
1
wrath815 :
Crawl but it’s progress
2026-09-15 01:02:44
0
Wren Ai :
its not that great, that and Polaris can run GLM 2.5 1T param but its like 0.3tps
2026-09-14 22:53:17
1
ZT :
Speed and execution = money which most of us can’t afford
2026-09-15 02:54:29
0
Robin Vanderbilt 📚✍️ Author :
Swap file have been around for years
2026-09-14 23:06:28
0
mastercho :
oh no we are not ready for ssd price to jump up to 80%
2026-09-14 20:37:41
3
benjamin :
sound awesome, won't it decrease the ssd drive life span with so meny writes?
2026-09-14 22:58:13
0
getklarted :
Definitely going to try it on my Mac mini M6
2026-09-14 20:23:12
1
Yuh-Roon (Jeroen) :
it's CO-li-bri, not co-LI-bri
2026-09-14 22:59:12
0
Mufaro :
1Tk / year
2026-09-15 01:08:18
0
LumaHug :
AI
2026-09-14 22:38:49
0
Helder Chaves :
ai
2026-09-14 22:20:36
0
AllThingsTruth1701 :
A build with enough concurrent channels to make this actually worth doing token per second-wise, would cost more than just building the rig in the first place 🤣 (probably). Also, at this point, anything can run anything.. the issue is output speed and 744B on 16gb no gpu laptop would be a week to respond, and 1 token output per day.. (dramatacising for effect). I'd be interested to know not if I can run a trillionB on my speak 'n' spell, but what's it's time to response and token output rates?
2026-09-14 21:50:46
0
__tarzan____ :
right it will be slow, but don't think with the ssd u have in mind right now, think few months later when that technology is fast enough
2026-09-14 22:07:39
0
🇳🇴 Cato 🇳🇴 :
AI
2026-09-14 22:05:20
0
lord pilkington :
Yeah but here's the thing: it's always going to be slow as shit. It's not an engineering problem, an optimization problem. It's physics. And then if people try to quantize stuff too much, the person that goes is going to be reasoning and math.
2026-09-14 22:42:19
0
ijpwivkqynk :
How does it work if I have a NVIDIA DGX Spark and want to run tier models could I use this to run Kimi or other like GLM 5.3?
2026-09-14 22:11:20
0
Ev1L Un1c0rN :
1 token per month?
2026-09-15 03:02:07
0
Asei Rola :
AI
2026-09-14 20:10:29
0
To see more videos from user @marcinteodoru, please go to the Tikwm
homepage.