-
Notifications
You must be signed in to change notification settings - Fork 34.6k
Add Unlimited OCR #46836
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Add Unlimited OCR #46836
Changes from all commits
13b0585
3d0a006
16a7260
ebf9b3b
a531dd8
8cd0c32
83ab5d8
b6a226e
e2d8d50
e5f2c46
0e47343
a26c74a
38f3199
6c42faa
1263664
b7c0276
34abc9f
6cf169b
0366fc3
73281a7
ec0c401
8ecbf2c
96a794a
d4aa878
bffafd0
16dcdbb
3bba66b
c81df8b
ab76c3a
8ed54f3
f608f97
a62fdc2
f346381
b91da12
4c3dc45
0270a2f
91e5996
e084630
15150c8
76307e9
07782c9
64ab0cd
0ec1eb0
0e9d3ad
5b40b2a
975c5b2
2d4d952
927be91
50f717a
390b786
a9eb8ae
c204543
a67620e
34c8d1d
3dffb4d
1d9a231
62883fc
37da184
95c7f1e
f37b18e
62065bc
bec037c
1ea87e3
033041e
f5b37fa
7b4db92
8a2a582
d91e66c
031cfd3
6b14e3d
d2a6e77
2d0505d
f2f5cab
c60d24c
e21c1d9
6dc6d88
c28f1da
0d85b89
3fddfb5
c8e1e2b
85b2e1b
1e8df96
21afbad
817810c
abdea09
f699734
d86f5c9
183fcf4
5024831
690cd92
9b54925
10737af
23aabee
b7f596f
ba1ed65
6b0b7d9
f2d05e9
7d5658f
7739c00
68a1f7a
3d77c49
7ab7f5c
81dc956
aa3f072
3fe8f6f
a729ebc
d5c0e01
fd4bb92
6cd1668
087b1df
55e6f54
ad5cdab
6e37e67
667666a
cb7bc6a
e1cfa15
66a63f1
0fdbf3c
17966e5
e388923
301622b
6b1fab5
ac7e12c
08eeed4
0acc0ce
93b9798
7dc2709
3c9ab38
7519723
a21362f
5236059
af7f77a
0ec84b6
95c34af
27de2f0
f0504a3
73c808f
ca8c3a7
8e0fbc9
7e14a2b
cc7d26e
bac08b3
5f37dfa
faf2f10
e39e73f
d8accb9
461a803
22c38d9
238bf3e
cc6edbb
d0636d6
f9ad01c
6ef423e
f690cd6
7196f34
a9c0772
5a96cc1
ca08ee3
4e74380
ccf90da
527a69d
0601965
6a81897
bb13eed
c47d816
2093858
40afe9d
9581dfc
8fe09cb
203e38c
f0fd46e
5a30585
6e31d53
7a2d9a9
36adb9f
ca024a3
5b6487f
c8d8626
977bc70
ea4d55a
fa90758
bf24d27
922231c
51afb7f
bb9341d
d4f5a8f
664f30b
3774c0e
96d4a29
3fdba03
745a3e2
1276563
3e1e537
94e7e40
417946d
6b5f625
6ffde03
df08460
f835baf
1657694
d7a8a18
0eefd0c
1eeb9ec
774d489
f620836
0ad2158
1cf342a
76076d7
5e21344
2065439
80d08e6
1a4109a
124ddef
5102060
f127652
1172fac
54b13d5
adb0e0e
6c0b38c
6a087ca
2788fbe
fb34ec1
1ab76a9
f81a69d
0680b54
30f2ff2
21f6dbf
9a0e9a7
c1cb494
c8cbb0d
67c57c6
f97ada9
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,264 @@ | ||
| <!--Copyright 2026 the HuggingFace Team. All rights reserved. | ||
|
|
||
| Licensed under the Apache License, Version 2.0 (the "License"); | ||
| you may not use this file except in compliance with the License. | ||
| You may obtain a copy of the License at | ||
|
|
||
| http://www.apache.org/licenses/LICENSE-2.0 | ||
|
|
||
| Unless required by applicable law or agreed to in writing, software | ||
| distributed under the License is distributed on an "AS IS" BASIS, | ||
| WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| See the License for the specific language governing permissions and | ||
| limitations under the License. | ||
|
|
||
|
|
||
| ⚠️ Note that this file is in Markdown but contain specific syntax for our doc-builder (similar to MDX) that may not be rendered properly in your Markdown viewer. | ||
|
|
||
| --> | ||
| *This model was published in HF papers on 2026-06-23 and contributed to Hugging Face Transformers on 2026-09-14.* | ||
|
|
||
|
|
||
| # UnlimitedOcr | ||
|
|
||
| ## Overview | ||
|
|
||
| The UnlimitedOcr model was proposed in [Unlimited OCR Works](https://huggingface.co/papers/2606.23050) by Youyang Yin, Huanhuan Liu, Qunyi Xie, Chaorun Liu, Shiqi Yang, Shaohua Wang, Zhanlong Liu, Hao Zou, Jinyue Chen, Shu Wei, Jingjing Wu, Mingxin Huang, Zhen Wu, Guibin Wang, Tengyu Du, and Lei Jia from Baidu Inc. | ||
|
|
||
| It is a single 3B parameter model with a standard context length of 32,768 tokens. | ||
|
|
||
| The abstract from the paper is the following: | ||
|
|
||
| *Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a large language model (LLM) as the decoder allows the model to leverage the prior distribution of language, leading to improved OCR performance. However, the downside is equally evident: as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. In this technical report, we propose Unlimited OCR, a model designed to emulate human parsing working memory. Taking DeepSeek OCR as the baseline, we replace all attention layers in the decoder with our proposed Reference Sliding Window Attention (R-SWA), which reduces attention computation costs while maintaining a constant KV cache throughout the entire decoding process. By combining the high compression rate of DeepSeek OCR's encoder with our constant KV cache design, Unlimited OCR can transcribe dozens of pages of documents in a single forward pass under a standard maximum length of 32K. More importantly, R-SWA is a general-purpose parsing attention mechanism — beyond OCR, it is equally applicable to tasks such as ASR, translation, etc. Codes and model weights are publicly available at http://github.com/baidu/Unlimited-OCR* | ||
|
|
||
| <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/unlimited_ocr_architecture.png" width="600"> | ||
|
|
||
| Unlimited-OCR supports two inference configurations: the default "gundam" mode uses 640x640 tiles with dynamic cropping for high-resolution documents, and "base" mode uses a single 1024x1024 global view for standard-resolution inputs. To enable base mode set `crop_to_patches=False` in the processor, or pass it via `processor_kwargs` when using [`~ProcessorMixin.apply_chat_template`]. | ||
|
|
||
| The vision tower follows the two-stage approach from [DeepSeek-OCR-2](./deepseek_ocr2): a SAM ViT-B encoder feeds into a CLIP ViT encoder. Unlike DeepSeek-OCR-2, the CLIP features are additionally concatenated with the SAM features to yield the final image tokens. Unlimited-OCR also omits the learnable patch queries from DeepSeek-OCR-2. | ||
|
|
||
| The text model is identical to DeepSeek-OCR-2 with the additional Reference Sliding Window Attention (R-SWA). R-SWA applies only to generated tokens. All image and prompt tokens remain fully visible throughout decoding, so long documents do not lose context from earlier pages. | ||
|
|
||
| This model was contributed by [guarin](https://huggingface.co/guarin). | ||
| The original code can be found [here](https://github.com/baidu/Unlimited-OCR). | ||
|
|
||
| > [!TIP] | ||
| > The original implementation runs the model with [torch.autocast](https://pytorch.org/docs/stable/amp.html#torch.autocast) in bfloat16. Wrap the forward and generate calls in an autocast context to reproduce its outputs. | ||
| > | ||
| > ```python | ||
| > with torch.autocast(device_type=model.device.type, dtype=torch.bfloat16): | ||
| > output = model.generate(**inputs) | ||
| > ``` | ||
|
|
||
| <hfoptions id="usage"> | ||
| <hfoption id="Single-page OCR"> | ||
|
|
||
| ```python | ||
| from transformers import AutoProcessor, AutoModelForImageTextToText | ||
|
|
||
| model = AutoModelForImageTextToText.from_pretrained("baidu/Unlimited-OCR", device_map="auto") | ||
| processor = AutoProcessor.from_pretrained("baidu/Unlimited-OCR") | ||
|
|
||
| image = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_suggestion_form.jpg" | ||
| messages = [ | ||
| { | ||
| "role": "user", | ||
| "content": [{"type": "image", "url": image}, {"type": "text", "text": "document parsing."}], | ||
| } | ||
| ] | ||
| inputs = processor.apply_chat_template( | ||
| messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt" | ||
| ).to(model.device, model.dtype) | ||
|
|
||
| output = model.generate(**inputs, max_new_tokens=512) | ||
| processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True) | ||
| # image [383, 87, 497, 171]\ntext [333, 201, 558, 230]R&D QUALITY IMPROVEMENT\nSUGGESTION/SOLUTION FORM... | ||
|
|
||
| # All bounding boxes are in (x1, y1, x2, y2) format with coordinates normalized to [0, 999] | ||
| ``` | ||
|
|
||
| ### Batch processing | ||
|
|
||
| For batch processing, pass a list of messages. Set `padding=True` for the processor if images have different sizes. | ||
|
|
||
| ```python | ||
| from transformers import AutoProcessor, AutoModelForImageTextToText | ||
|
|
||
| model = AutoModelForImageTextToText.from_pretrained("baidu/Unlimited-OCR", device_map="auto") | ||
| processor = AutoProcessor.from_pretrained("baidu/Unlimited-OCR") | ||
|
|
||
| image1 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_suggestion_form.jpg" | ||
| image2 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_receipt.jpeg" | ||
| messages = [ | ||
| [ | ||
| { | ||
| "role": "user", | ||
| "content": [{"type": "image", "url": image}, {"type": "text", "text": "document parsing."}], | ||
| } | ||
| ] | ||
| for image in [image1, image2] | ||
| ] | ||
| inputs = processor.apply_chat_template( | ||
| messages, | ||
| add_generation_prompt=True, | ||
| tokenize=True, | ||
| return_dict=True, | ||
| return_tensors="pt", | ||
| processor_kwargs={"padding": True}, | ||
| ).to(model.device, model.dtype) | ||
|
|
||
| output = model.generate(**inputs) | ||
| processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True) | ||
| # image [383, 87, 497, 171]\ntext [333, 201, 558, 230]R&D QUALITY IMPROVEMENT\nSUGGESTION/SOLUTION FORM... | ||
|
|
||
| # All bounding boxes are in (x1, y1, x2, y2) format with coordinates normalized to [0, 999] | ||
| ``` | ||
|
|
||
| ### Region detections | ||
|
|
||
| Set `skip_special_tokens=False` to wrap all detections and region types in `<|det|>...<|/det|>` markers. This is useful for further post-processing of the output, for example to plot the detected bounding boxes on the image. Each detection is wrapped as `<|det|>region_type [x1, y1, x2, y2]<|/det|>text...` with coordinates normalized to a `[0, 999]` range. Set `return_detections=True` to get an additional list of dictionaries with all detections parsed as `{"region_type": region_type, "box": [x1, y1, x2, y2], "text": "..."}`. | ||
|
|
||
| ```python | ||
| from transformers import AutoProcessor, AutoModelForImageTextToText | ||
|
|
||
| model = AutoModelForImageTextToText.from_pretrained("baidu/Unlimited-OCR", device_map="auto") | ||
| processor = AutoProcessor.from_pretrained("baidu/Unlimited-OCR") | ||
|
|
||
| image = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_suggestion_form.jpg" | ||
| messages = [ | ||
| { | ||
| "role": "user", | ||
| "content": [{"type": "image", "url": image}, {"type": "text", "text": "document parsing."}], | ||
| } | ||
| ] | ||
| inputs = processor.apply_chat_template( | ||
| messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt" | ||
| ).to(model.device, model.dtype) | ||
|
|
||
| output = model.generate(**inputs) | ||
| decoded, detections = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=False, return_detections=True) | ||
| # <|det|>image [383, 87, 497, 171]<|/det|>\n<|det|>text [333, 201, 558, 230]<|/det|>R&D QUALITY IMPROVEMENT\nSUGGESTION/SOLUTION FORM... | ||
|
|
||
| # Visualization | ||
| import random | ||
| import matplotlib.pyplot as plt | ||
| import matplotlib.patches as patches | ||
| from transformers.image_utils import load_image | ||
|
|
||
| image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_suggestion_form.jpg") | ||
| width, height = image.size | ||
|
|
||
| figure, axis = plt.subplots(figsize=(10, 12)) | ||
| axis.imshow(image) | ||
| for detection in detections: | ||
| region_type = detection["region_type"] | ||
| x1, y1, x2, y2 = detection["box"] | ||
| x1, y1, x2, y2 = int(x1) / 999 * width, int(y1) / 999 * height, int(x2) / 999 * width, int(y2) / 999 * height | ||
| color = (random.random(), random.random(), random.random()) | ||
| rectangle = patches.Rectangle((x1, y1), x2 - x1, y2 - y1, linewidth=1.5, edgecolor=color, facecolor="none") | ||
| axis.add_patch(rectangle) | ||
| axis.text(x1, y1, region_type, color="white", fontsize=8, backgroundcolor=color, verticalalignment="top") | ||
| axis.axis("off") | ||
| plt.show() | ||
| ``` | ||
|
|
||
| <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/unlimited_ocr_suggestion_form_boxes.jpg" width="600"> | ||
|
|
||
| </hfoption> | ||
| <hfoption id="Multi-page OCR"> | ||
|
|
||
| Multi-page documents can be parsed jointly in a single forward pass by passing all page images together. Add one image block per page to the message so the model processes all pages as a continuous document. | ||
|
|
||
| ```python | ||
| from transformers import AutoProcessor, AutoModelForImageTextToText | ||
|
|
||
| model = AutoModelForImageTextToText.from_pretrained("baidu/Unlimited-OCR", device_map="auto") | ||
| processor = AutoProcessor.from_pretrained("baidu/Unlimited-OCR") | ||
|
|
||
| page1 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_suggestion_form.jpg" | ||
| page2 = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/ocr_receipt.jpeg" | ||
|
|
||
| messages = [ | ||
| { | ||
| "role": "user", | ||
| "content": [{"type": "image", "url": page} for page in [page1, page2]] | ||
| + [{"type": "text", "text": "Multi page parsing."}], | ||
| } | ||
| ] | ||
| inputs = processor.apply_chat_template( | ||
| messages, | ||
| add_generation_prompt=True, | ||
| tokenize=True, | ||
| return_dict=True, | ||
| return_tensors="pt", | ||
| processor_kwargs={"crop_to_patches": False}, | ||
| ).to(model.device, model.dtype) | ||
|
|
||
| output = model.generate( | ||
| **inputs, | ||
| no_repeat_ngram_window_size=1024, # default: 128, larger window size recommended for long documents | ||
| ) | ||
| processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True) | ||
| # <PAGE>image [382, 87, 489, 174]\ntitle [333, 201, 556, 230]R&D QUALITY IMPROVEMENT\nSUGGESTION/SCLUTION FORM... | ||
|
|
||
| # All bounding boxes are in (x1, y1, x2, y2) format with coordinates normalized to [0, 999] | ||
| ``` | ||
|
|
||
| </hfoption> | ||
| </hfoptions> | ||
|
|
||
| ## UnlimitedOcrConfig | ||
|
|
||
| [[autodoc]] UnlimitedOcrConfig | ||
|
|
||
| ## UnlimitedOcrTextConfig | ||
|
|
||
| [[autodoc]] UnlimitedOcrTextConfig | ||
|
|
||
| ## UnlimitedOcrVisionConfig | ||
|
|
||
| [[autodoc]] UnlimitedOcrVisionConfig | ||
|
|
||
| ## UnlimitedOcrVisionEncoderConfig | ||
|
|
||
| [[autodoc]] UnlimitedOcrVisionEncoderConfig | ||
|
|
||
| ## UnlimitedOcrSamVisionConfig | ||
|
|
||
| [[autodoc]] UnlimitedOcrSamVisionConfig | ||
|
|
||
| ## UnlimitedOcrImageProcessor | ||
|
|
||
| [[autodoc]] UnlimitedOcrImageProcessor | ||
|
|
||
| ## UnlimitedOcrProcessor | ||
|
|
||
| [[autodoc]] UnlimitedOcrProcessor | ||
|
|
||
| ## UnlimitedOcrPreTrainedModel | ||
|
|
||
| [[autodoc]] UnlimitedOcrPreTrainedModel | ||
|
|
||
| ## UnlimitedOcrTextPreTrainedModel | ||
|
|
||
| [[autodoc]] UnlimitedOcrTextPreTrainedModel | ||
|
|
||
| ## UnlimitedOcrTextModel | ||
|
|
||
| [[autodoc]] UnlimitedOcrTextModel | ||
| - forward | ||
|
|
||
| ## UnlimitedOcrVisionModel | ||
|
|
||
| [[autodoc]] UnlimitedOcrVisionModel | ||
| - forward | ||
|
|
||
| ## UnlimitedOcrModel | ||
|
|
||
| [[autodoc]] UnlimitedOcrModel | ||
| - forward | ||
|
|
||
| ## UnlimitedOcrForConditionalGeneration | ||
|
|
||
| [[autodoc]] UnlimitedOcrForConditionalGeneration | ||
| - forward |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -163,7 +163,8 @@ | |
| "image_sizes_videos", | ||
| "pixel_attention_mask", | ||
| "pixel_values_images", | ||
| "num_local_patches", | ||
|
guarin marked this conversation as resolved.
|
||
| "pixel_values_local", | ||
| "patches_grid", | ||
|
Collaborator
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. potentially different PR (to first cover the deepseek model) or is it only now with unlimited ocr?
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
|
||
| ) | ||
|
|
||
|
|
||
|
|
||
Uh oh!
There was an error while loading. Please reload this page.