@rick.theengineer: Every answer from ChatGPT or Claude is a fight over 80 GB. 🧠⚡ Training a model happens once. Serving it to millions of people, every second of every day, is the part nobody sees. Here's what actually happens after you hit send: 1️⃣ The model doesn't fit. A 100-billion-parameter model is about 200 GB at 16-bit. One NVIDIA H100 holds 80 GB. So the model gets split across several GPUs, and they have to talk over NVLink for every single token. 2️⃣ Your prompt is read all at once (prefill). The answer is not. Every word you see is predicted one token at a time. Token 100 literally cannot exist before token 99. 3️⃣ One GPU, dozens of people. Inference servers like vLLM batch users together, and every step writes one token for everyone in the batch. When someone finishes, a new user takes their lane on the very next step. That scheduling alone can double what the same GPU gets done. 4️⃣ The hidden memory hog: the KV cache. The model keeps notes on every token of your conversation. For Llama 3.1 70B, one full 128K-token conversation needs about 43 GB of cache. That's over half a GPU for one chat. 5️⃣ The tricks that make it affordable: → quantization (16 bits down to 8 or 4) → tensor + pipeline parallelism → mixture of experts (Mixtral holds 46.7B parameters but uses 12.9B per token) 6️⃣ Why words stream in. Tokens are sent the moment they're generated, so a 20-second answer feels instant. What matters: time to first token, and how fast the rest follow. Training builds the brain. Inference is the infrastructure problem of letting millions of people use it at once without spending a fortune every second. 💬 Which part surprised you most? Drop a number 1–6 👇 📌 Save this for the next time someone says AI is "just an API call." #aiinference #llm #machinelearning #nvidia #softwareengineering
RickTheEngineer
Region: GB
Saturday 26 September 2026 17:41:56 GMT
Music
Download
Comments
Trexaki :
Amazing Illustration and without over explanation.[Μου αρέσει]
2026-09-27 00:34:55
1
Alex12345 :
can u explain slowly because my brain is slow
2026-09-27 11:13:43
0
Csaba :
so nothing new....👍
2026-09-26 22:27:13
1
Inspector Mango Bango :
thanks rick
2026-09-26 19:41:44
1
To see more videos from user @rick.theengineer, please go to the Tikwm
homepage.