@qoima_qsm: Правиланы бұзыпты ғой 😱 #qoima

qoima_qsm
qoima_qsm
Open In TikTok:
Region: KZ
Thursday 10 September 2026 16:25:09 GMT
97000
2516
131
74

Music

Download

Comments

nurperi1071
Нурпери 🇰🇬 :
даусы жок
2026-09-10 17:01:48
322
akosya7878
yoru :
даусы барго немене жок деп атсндар
2026-09-11 04:37:57
76
arai.polat8
Arai Polat :
даусы какта
2026-09-30 05:14:38
2
aluka7714
~Aluka~ :
звук
2026-09-11 07:19:32
25
077b08
𝓑𝓲𝓴𝓸シ︎ :
қайттан салыңыздарш даусы жоұ
2026-09-23 07:37:03
0
aiizh65
Aiizh :
даусы қайдеей халкым
2026-10-01 04:47:38
3
aisha_baiansulu_
Aisha🖤 :
менде ғана дауысы жоқ па😳
2026-09-13 07:48:32
18
gulnazik640
гули🌺❤️✌️ :
не даусы кайда мен гана ма не
2026-09-11 13:25:29
32
zhannurrr67
𝔃𝓱𝓪𝓷𝓷𝓾𝓻 :
субриты….
2026-09-13 05:22:22
4
fanzanuli43
ФАН ЖАНУЛИ :
дауыз жок менде гана ма
2026-09-15 06:11:15
1
erke_goi097
erke_goi0 :
дауысықайда😳
2026-09-13 08:22:21
2
andrygarfilb
садр фан :
даусы қайдаааааа
2026-09-20 10:48:48
3
boks_3102
D🪽 :
дауска не болд
2026-09-20 04:31:07
2
kamoooo0000
𝖐 :
Дауыс ғде?
2026-09-12 04:18:42
19
aimyra331
Akoaboo)) :
даусы гдесн
2026-09-10 18:39:18
3
madina.03_
Madina :
ун жок
2026-09-12 04:15:45
4
aibi5727
А. :
Даусы какта
2026-09-11 05:54:42
10
inzhuk511
Инжукаа🎀 :
Субтитр окып откан мен 😁
2026-09-12 19:47:38
5
zere2910
M.Batyrkhankyzy❤️🌹 :
неге дауысы жок
2026-09-15 03:02:36
1
aiym_goi5
аkoo🪼 :
1
2026-09-10 16:29:16
2
assem_abdrahmanova
𝑨𝒔𝒔𝒆𝒎𝑺𝒐𝒔𝒌𝒂𝒂 :
Дауыс
2026-09-10 18:42:45
1
assem_abdrahmanova
𝑨𝒔𝒔𝒆𝒎𝑺𝒐𝒔𝒌𝒂𝒂 :
Даусы
2026-09-10 18:42:41
1
013.tkd.almat
ткд💋🔥🥇 :
дауыс
2026-09-10 17:31:40
1
maratova_770
Maratova_G :
дауысы бар гой
2026-09-11 13:52:13
2
amange1d1vna
Nargiza :
Касында адамы неге айтпайт шартты бузды деп🤦🏻‍♀️
2026-09-12 06:58:00
2
To see more videos from user @qoima_qsm, please go to the Tikwm homepage.

Other Videos

What exactly is model distillation — and why did DeepSeek suddenly make everyone talk about it? Think of it as a teacher and a student. The teacher is a massive model: very smart, very expensive, very slow. The student is a much smaller one: cheap and fast, but not as capable. Instead of training the small model only on raw internet data, you let the big model write the answers — the worked examples, the reasoning, step by step. Those outputs become the student's training data. The student never copies the teacher's weights. It learns the pattern. The result isn't as smart as the teacher. But it's dramatically cheaper and faster to run — and surprisingly strong. DeepSeek did exactly this with R1: they used the big model to generate reasoning data, then trained smaller Qwen and Llama models on it, from 1.5B all the way to 70B. Their own report found that distilling from R1 worked better than making those small models discover the reasoning themselves through reinforcement learning. Then the controversy. OpenAI has accused DeepSeek of using outputs from OpenAI models — accounts, programmatic access, bypassed restrictions. DeepSeek hasn't confirmed it. Two things to keep apart: ✅ DeepSeek distilling its own R1 into smaller models — documented. ⚠️ DeepSeek distilling from OpenAI models — an allegation, not an established fact. And distillation itself is not shady. It's a standard machine learning method. DeepSeek's own MIT licence explicitly allows using R1 outputs for fine-tuning and distillation. The real question is permission. If the model owner allows it, distillation is completely normal. If the terms prohibit building a competing model, doing it anyway is a very different issue. The one-line version: the biggest model discovers the capability. Distillation squeezes as much of it as possible into something smaller. Save this for the next time someone says
What exactly is model distillation — and why did DeepSeek suddenly make everyone talk about it? Think of it as a teacher and a student. The teacher is a massive model: very smart, very expensive, very slow. The student is a much smaller one: cheap and fast, but not as capable. Instead of training the small model only on raw internet data, you let the big model write the answers — the worked examples, the reasoning, step by step. Those outputs become the student's training data. The student never copies the teacher's weights. It learns the pattern. The result isn't as smart as the teacher. But it's dramatically cheaper and faster to run — and surprisingly strong. DeepSeek did exactly this with R1: they used the big model to generate reasoning data, then trained smaller Qwen and Llama models on it, from 1.5B all the way to 70B. Their own report found that distilling from R1 worked better than making those small models discover the reasoning themselves through reinforcement learning. Then the controversy. OpenAI has accused DeepSeek of using outputs from OpenAI models — accounts, programmatic access, bypassed restrictions. DeepSeek hasn't confirmed it. Two things to keep apart: ✅ DeepSeek distilling its own R1 into smaller models — documented. ⚠️ DeepSeek distilling from OpenAI models — an allegation, not an established fact. And distillation itself is not shady. It's a standard machine learning method. DeepSeek's own MIT licence explicitly allows using R1 outputs for fine-tuning and distillation. The real question is permission. If the model owner allows it, distillation is completely normal. If the terms prohibit building a competing model, doing it anyway is a very different issue. The one-line version: the biggest model discovers the capability. Distillation squeezes as much of it as possible into something smaller. Save this for the next time someone says "distilled model" like you're supposed to know what it means. #modeldistillation #deepseek #machinelearning #aiexplained #llm

About