@dianasaurbytes: **DeepSeek-OCR:** A novel vision encoder architecture for compressing text into visual representations while achieving high OCR accuracy. - [GitHub Repository](https://github.com/deepseek-ai/DeepSeek-OCR) - [Research Paper](https://www.arxiv.org/pdf/2510.18234) Let’s discuss DeepSeek-OCR, an innovative open-source model by Deepseek for optical character recognition (OCR). This paper presents an exciting approach to using images to encode textual information more efficiently. Instead of converting long documents into thousands of tokens for LLMs, **DeepSeek‑OCR converts entire pages into images**, which are then compressed into a smaller set of “vision tokens.” This method results in 7–20x fewer tokens compared to traditional text representation and maintains impressive accuracy—97% at 10x compression and about 60% at 20x. ### Key Components: 1. **DeepEncoder:** - Combines: - **SAM-base** (Segment Anything) — local attention for perception. - **CLIP-large** — global semantic understanding. - Incorporates a **16× convolutional token compressor** that fuses these components to extract features, tokenize them, and compress the information. 2. **Decoder:** - Named Deepseek-3B-MoE, this transformer-based model takes the vision tokens and decodes them back into text. This suggests a new paradigm where large documents may be stored as compressed visual representations rather than lengthy token sequences—a development that could significantly enhance compute efficiency, reduce memory usage, and lower latency when processing multi-page documents like PDFs or research papers. #DeepLearning #ComputerVision #AIResearch #Innovation #Tokenization #VisualRepresentation #MachineLearningModel #OpenSourceAI #OCRTechnology #EfficiencyInComputing
dianasaurbytes
Region: US
Sunday 07 December 2025 02:15:38 GMT
Music
Download
Comments
appleuser69959680 :
So, a picture really is worth a thousand words 👍
2025-12-07 22:27:26
2
Blue_Robin :
So how do you use it?
2025-12-07 16:47:10
0
UrbanGringo :
this is a similar model to minicpm model. my agent uses it to understand what she's looking at, so the information is a bit old old on this front
2026-06-01 15:39:48
0
AlamIndah :
this is an older paper right?
2025-12-07 03:44:34
1
X :
Isn’t SAM an open source model from Meta?
2025-12-09 12:42:57
0
Salvatore :
so for text we have embeddings that we store in a vector base what about vision tokens where do we store them and how do we make sense of them and how we get them ?
2026-01-10 19:35:35
0
Juan Erke :
Nice explanation. thank you 👍
2025-12-09 11:51:06
0
datasyed :
Great
2025-12-10 14:30:12
0
Fish369 :
How do you stop hallucinations?
2025-12-10 20:31:47
0
Binary_Harmony :
is there any image recognition and text recognitionembedded in ocr tech right now, like reasoning of logos image/ graph to text utilization in pdf ?
2025-12-12 14:54:34
0
Bloody Kheeng :
php imagik extension with tessaract ocr is better 🤔
2025-12-08 16:43:12
0
لمى :
Next time please explain it in 5 words. I suffer from brain rot 😭
2026-02-02 22:12:54
0
D12 :
🥰🥰🥰
2025-12-17 16:30:47
0
naz :
@Habib | DS enthusiast
2025-12-18 12:01:01
0
1มกราคม#ป้ายหาดใหญ่ :
🥰🥰
2026-04-28 19:03:48
0
Numbers [Dead Dove Do Not Eat] :
OCR such trash, I’m always wanting to jump off of a bridge when I have to leverage it
2025-12-29 18:09:44
0
George Robertson's Music :
What’s the api
2025-12-23 20:47:14
0
To see more videos from user @dianasaurbytes, please go to the Tikwm
homepage.