
DeepSeek OCR introduces a novel approach to artificial intelligence memory by implementing "artificial forgetting" through high-ratio data compression. By rendering text and documents as images and applying learned, token-based compression, the system reduces data size by up to 90% while maintaining over 96% text recovery accuracy. This efficiency is achieved through a specialized encoder utilizing local window attention, a 16x convolutional compressor, and global visual processing, allowing the model to skim and squeeze high-resolution documents without prohibitive computational costs. Beyond document recognition, the system generates training data for large language models at speeds exceeding 200,000 pages per day on a single GPU. While the model occasionally relies on language priors to "guess" highly compressed text, it represents a significant shift toward more sustainable, open-source AI development by drastically lowering the cost of processing and remembering vast amounts of information.
Sign in to continue reading, translating and more.
Open full episode in Podwise