@dianasaurbytes: **DeepSeek-OCR:** A novel vision encoder architecture for compressing text into visual representations while achieving high OCR accuracy. - [GitHub Repository](https://github.com/deepseek-ai/DeepSeek-OCR) - [Research Paper](https://www.arxiv.org/pdf/2510.18234) Let’s discuss DeepSeek-OCR, an innovative open-source model by Deepseek for optical character recognition (OCR). This paper presents an exciting approach to using images to encode textual information more efficiently. Instead of converting long documents into thousands of tokens for LLMs, **DeepSeek‑OCR converts entire pages into images**, which are then compressed into a smaller set of “vision tokens.” This method results in 7–20x fewer tokens compared to traditional text representation and maintains impressive accuracy—97% at 10x compression and about 60% at 20x. ### Key Components: 1. **DeepEncoder:** - Combines: - **SAM-base** (Segment Anything) — local attention for perception. - **CLIP-large** — global semantic understanding. - Incorporates a **16× convolutional token compressor** that fuses these components to extract features, tokenize them, and compress the information. 2. **Decoder:** - Named Deepseek-3B-MoE, this transformer-based model takes the vision tokens and decodes them back into text. This suggests a new paradigm where large documents may be stored as compressed visual representations rather than lengthy token sequences—a development that could significantly enhance compute efficiency, reduce memory usage, and lower latency when processing multi-page documents like PDFs or research papers. #DeepLearning #ComputerVision #AIResearch #Innovation #Tokenization #VisualRepresentation #MachineLearningModel #OpenSourceAI #OCRTechnology #EfficiencyInComputing

dianasaurbytes
dianasaurbytes
Open In TikTok:
Region: US
Sunday 07 December 2025 02:15:38 GMT
29315
1256
22
182

Music

Download

Comments

appleuser69959680
appleuser69959680 :
So, a picture really is worth a thousand words 👍
2025-12-07 22:27:26
2
blue_robin_usb
Blue_Robin :
So how do you use it?
2025-12-07 16:47:10
0
urbangringo6
UrbanGringo :
this is a similar model to minicpm model. my agent uses it to understand what she's looking at, so the information is a bit old old on this front
2026-06-01 15:39:48
0
tokoalamindah
AlamIndah :
this is an older paper right?
2025-12-07 03:44:34
1
xarain8888
X :
Isn’t SAM an open source model from Meta?
2025-12-09 12:42:57
0
salvatore_0_2
Salvatore :
so for text we have embeddings that we store in a vector base what about vision tokens where do we store them and how do we make sense of them and how we get them ?
2026-01-10 19:35:35
0
juanerke
Juan Erke :
Nice explanation. thank you 👍
2025-12-09 11:51:06
0
datasyed
datasyed :
Great
2025-12-10 14:30:12
0
dawsamfl
Fish369 :
How do you stop hallucinations?
2025-12-10 20:31:47
0
asaf.avidan
Binary_Harmony :
is there any image recognition and text recognitionembedded in ocr tech right now, like reasoning of logos image/ graph to text utilization in pdf ?
2025-12-12 14:54:34
0
bloodykheeng
Bloody Kheeng :
php imagik extension with tessaract ocr is better 🤔
2025-12-08 16:43:12
0
yhcdruhd467
لمى :
Next time please explain it in 5 words. I suffer from brain rot 😭
2026-02-02 22:12:54
0
d12_city
D12 :
🥰🥰🥰
2025-12-17 16:30:47
0
nrppppopp
naz :
@Habib | DS enthusiast
2025-12-18 12:01:01
0
user91784085586048
1มกราคม#ป้ายหาดใหญ่ :
🥰🥰
2026-04-28 19:03:48
0
numberstheperson
Numbers [Dead Dove Do Not Eat] :
OCR such trash, I’m always wanting to jump off of a bridge when I have to leverage it
2025-12-29 18:09:44
0
_george_robertson_
George Robertson's Music :
What’s the api
2025-12-23 20:47:14
0
To see more videos from user @dianasaurbytes, please go to the Tikwm homepage.

Other Videos


About