SEERRAY LAB / 0.8B OCR MODEL

Xiaomi-OCR-0

Read. Parse. Understand.

Document parsing and OCR-centric understanding, together in a compact 0.8B vision-language model.

A compact OCR reader recovering text, tables, and formulas from a document.
95.24Real5-OmniDocBenchOverall across five conditions 96.83OmniDocBench v1.6Overall parsing score 87.94Wild-OmniDocBenchPhysical-scene document parsing 83.2OCR-oriented VQAMean across five benchmarks ~170MOCR-centric samplesReported training corpus

01 / CAPABILITIES

Supported tasks.

[ 01 ]

Document parsing

[ 02 ]

Visual question answering

[ 03 ]

Key information extraction

SAMPLE PLAYBACK

Explore model outputs.

Choose a document, KIE, or VQA case to replay its saved result.

RECORDED OUTPUT · TYPEWRITER REPLAY
INPUT—

Choose a sample

Select one from the gallery to begin.

INPUT

The saved prompt will appear here.

MODEL OUTPUTSAVED RESULT

Select a case to play its saved response.

RECORDED SAMPLE

02 / BENCHMARKS

Evaluation results.

Document parsing and visual question answering. Select a benchmark to view the reported scores.

ACQUISITION ROBUSTNESS

Real5-OmniDocBench

95.24

XIAOMI-OCR-0 / OVERALL SCORE

Overall score and results for five document acquisition conditions.

01

Parameter efficiency

Score at a given model size

↖ Fewer parameters · higher score889092949600.30.60.91.21.5Model parameters (B) · smaller ←Overall score · higher →PaddleOCR-VL-1.6 · 0.9B · 93.19OvisOCR2 · 0.8B · 92.29PaddleOCR-VL-1.5 · 0.9B · 92.05GLM-OCR · 0.9B · 90.32MinerU2.5-Pro · 1.2B · 88.94Xiaomi-OCR-0 · 0.8B · 95.24Xiaomi-OCR-00.8B · 95.24
★ Xiaomi-OCR-0Colors & shapes identify models

Hover or select a point to view its reported scores.

02

Score comparison

Overall parsing score · higher is better

MODEL DETAILS
Xiaomi-OCR-0

0.8B parameters

95.24
Scanning
96.41
Warping
95.46
Screen photo
94.47
Illumination
95.39
Skew
94.46
View detailed scores
Selected comparisons from the technical report. Column arrows indicate the preferred direction.
ModelSizeOverallScanningWarpingScreen photoIlluminationSkew
Xiaomi-OCR-00.8B95.2496.4195.4694.4795.3994.46
PaddleOCR-VL-1.60.9B93.1994.7492.4892.7893.2892.66
OvisOCR20.8B92.2993.7791.4093.0992.8890.33
PaddleOCR-VL-1.50.9B92.0593.4391.2591.7692.1691.66
GLM-OCR0.9B90.3292.6790.6891.7591.1285.39
Kimi-K2.61T89.7690.0889.6289.5889.9189.61
Gemini-3 Pro-89.2489.4788.9088.8689.5389.45
MonkeyOCRv2-B-Parsing0.7B89.2289.4989.7088.4088.5489.97
Kimi K2.51T89.0989.6788.8688.3989.6688.86
Doubao-Seed-2.1-Pro-89.0288.8589.3688.9989.1388.79
MinerU2.5-Pro1.2B88.9492.1188.7291.2991.3181.26
Qwen3-VL-235B235B88.9089.4389.9989.2789.2786.56
Gemini-2.5 Pro-88.2189.2587.6387.1187.9789.07
MonkeyOCRv2-S-Parsing0.6B87.9088.8788.1787.6486.7588.09
Qwen2.5-VL-72B72B86.9286.1987.7786.4887.2586.90
dots.ocr3B86.3886.8786.0187.1887.5784.27
MinerU2.51.2B85.6190.0683.7689.4189.5775.24
PaddleOCR-VL0.9B85.5492.1185.9782.5489.6177.47
Nanonets-OCR-s3B84.1985.5283.5684.8685.0181.98
MonkeyOCR-pro-3B3.7B79.4986.9478.9082.4484.7164.47
GPT-5.2-78.6684.4376.2676.7580.8875.00
MonkeyOCR-3B3.7B78.2984.6577.2780.7183.1665.67
MonkeyOCR-pro-1.2B1.9B77.1584.6476.5980.2482.1162.18
MinerU2-VLM0.9B76.9583.6073.7378.7780.5168.16
DeepSeek-OCR3B73.9986.1767.2075.3178.1063.01
DeepSeek-OCR 23B73.0189.5966.5371.6576.0261.28
PP-StructureV3--64.4584.6859.3466.8973.3837.98
Dolphin322M61.7872.1660.3564.2967.2944.83
Dolphin-1.50.3B61.4883.3950.5069.7675.6128.16
Marker-1.8.2--60.1070.2758.9863.6566.3141.27

DOCUMENT PARSING

OmniDocBench v1.6

96.83

XIAOMI-OCR-0 / OVERALL SCORE

Document parsing scores, including text, tables, formulas, and reading order.

01

Parameter efficiency

Score at a given model size

↖ Fewer parameters · higher score889092949600.511.522.533.5Model parameters (B) · smaller ←Overall score · higher →TeleOCR · 1.2B · 96.87OvisOCR2 · 0.8B · 96.58PaddleOCR-VL-1.6 · 0.9B · 96.33MinerU2.5-Pro · 1.2B · 95.75GLM-OCR · 0.9B · 95.22PaddleOCR-VL-1.5 · 0.9B · 94.93HunyuanOCR-1.5 · 1B · 94.74PaddleOCR-VL · 0.9B · 94.18Unlimited-OCR · 3B-A0.5B · 93.92Youtu-Parsing · 2.5B · 93.74ABot-OCR · 2B · 93.30FireRed-OCR · 2B · 93.26dots.ocr · 3B · 90.77OpenDoc-0.1B · 0.1B · 90.67DeepSeek-OCR-2 · 3B · 90.25Dolphin-v2 · 3B · 89.50MonkeyOCR-pro-3B · 3B · 88.57Xiaomi-OCR-0 · 0.8B · 96.83Xiaomi-OCR-00.8B · 96.83
★ Xiaomi-OCR-0Colors & shapes identify models

Hover or select a point to view its reported scores. Select overlapping models from the score list.

Chart range: 0–3.5B · scores ≥88. All results remain in the score list.SVG ↗ PNG ↗
02

Score comparison

Overall parsing score · higher is better

MODEL DETAILS
Xiaomi-OCR-0

0.8B parameters

96.83
Text Edit ↓
0.031
Formula CDM ↑
98.50
Table TEDS ↑
95.11
Reading Order Edit ↓
0.122
View detailed scores
Selected comparisons from the technical report. Column arrows indicate the preferred direction.
ModelSizeOverallText Edit ↓Formula CDM ↑Table TEDS ↑Table TEDS-S ↑Reading Order Edit ↓
Xiaomi-OCR-00.8B96.830.03198.5095.1197.190.122
TeleOCR1.2B96.870.02796.3697.0598.520.122
OvisOCR20.8B96.580.02597.5394.7697.160.111
PaddleOCR-VL-1.60.9B96.330.03397.4994.7697.110.127
MinerU2.5-Pro1.2B95.750.03697.4593.4295.920.120
GLM-OCR0.9B95.220.04497.1892.8395.390.133
PaddleOCR-VL-1.50.9B94.930.03896.8991.6794.370.130
HunyuanOCR-1.51B94.740.03994.5093.6794.710.129
PaddleOCR-VL0.9B94.180.04095.9190.6593.740.135
Unlimited-OCR3B-A0.5B93.920.04295.7990.1693.320.129
Qianfan-OCR4B93.900.04095.0890.5393.310.130
Youtu-Parsing2.5B93.740.04493.6392.0295.000.116
Logics-Parsing-v24B93.330.04195.6588.4291.980.137
ABot-OCR2B93.300.03794.8688.6991.870.137
FireRed-OCR2B93.260.03795.4488.0491.060.131
dots.ocr3B90.770.04889.9587.1890.580.138
OpenDoc-0.1B0.1B90.670.04993.0283.8887.450.140
DeepSeek-OCR-23B90.250.05091.8483.8987.750.144
Dolphin-v23B89.500.06991.0184.4087.440.150
OCRVerse4B88.600.06389.6182.4486.270.163
MonkeyOCR-pro-3B3B88.570.07488.7484.3588.620.189
Dolphin-1.50.3B86.520.09487.4981.4384.820.167
olmOCR7B85.740.13988.1083.0087.170.216
Nanonets-OCR-s3B83.610.10881.4680.1884.510.213
POINTS-Reader3B83.370.09685.7273.9877.400.198

PHYSICAL-SCENE ROBUSTNESS

Wild-OmniDocBench

87.94

XIAOMI-OCR-0 / OVERALL SCORE

Xiaomi-OCR-0 scores 87.94 with 0.8B parameters on physically recaptured documents. Its Reading Order Edit of 0.1931 is the lowest among the methods compared in the report; TeleOCR leads overall with 88.53.

01

Parameter efficiency

Score at a given model size

↖ Fewer parameters · higher score75788184879001234Model parameters (B) · smaller ←Overall score · higher →TeleOCR · 1.2B · 88.53OvisOCR2 · 0.8B · 87.91PaddleOCR-VL-1.6 · 0.9B · 87.36MinerU2.5-Pro · 1.2B · 87.33GLM-OCR · 0.9B · 85.08PaddleOCR-VL-1.5 · 0.9B · 84.64dots.ocr · 3B · 81.84HunyuanOCR-1.5 · 1B · 77.62Logics-Parsing-v2 · 4B · 77.10Xiaomi-OCR-0 · 0.8B · 87.94Xiaomi-OCR-00.8B · 87.94
★ Xiaomi-OCR-0Colors & shapes identify models

Hover or select a point to view its reported scores. Select overlapping models from the score list.

02

Score comparison

Overall parsing score · higher is better

MODEL DETAILS
Xiaomi-OCR-0

0.8B parameters

87.94
Text Edit ↓
0.1233
Formula CDM ↑
89.55
Table TEDS ↑
86.61
Reading Order Edit ↓
0.1931
View detailed scores
Selected comparisons from the technical report. Column arrows indicate the preferred direction.
ModelSizeOverallText Edit ↓Formula CDM ↑Table TEDS ↑Table TEDS-S ↑Reading Order Edit ↓
Xiaomi-OCR-00.8B87.940.123389.5586.6190.940.1931
TeleOCR1.2B88.530.117388.2689.0592.140.2011
OvisOCR20.8B87.910.129090.3785.1389.110.2021
PaddleOCR-VL-1.60.9B87.360.136988.4285.7690.140.2057
MinerU2.5-Pro1.2B87.330.136290.1585.4690.120.2013
GLM-OCR0.9B85.080.151489.0981.3185.900.2228
PaddleOCR-VL-1.50.9B84.640.146186.7281.8086.520.2138
dots.ocr3B81.840.148385.0075.3280.200.2200
HunyuanOCR-1.51B77.620.197985.1267.5470.670.2750
Logics-Parsing-v24B77.100.402991.4080.1987.160.2355

VISUAL QUESTION ANSWERING

OCR-oriented VQA

83.2

XIAOMI-OCR-0 / 5-BENCHMARK MEAN

Xiaomi-OCR-0 scores 83.2 with 0.8B parameters, above every specialized OCR model in this comparison. It also exceeds Qwen3.5-2B (80.9) and MiniCPM-V-4.5 (82.6). Qwen3.5-4B scores 84.9 with five times as many parameters.

01

Parameter efficiency

Score at a given model size

↖ Fewer parameters · higher score56606468727680848802468Model parameters (B) · smaller ←Overall score · higher →Qwen3.5-4B · 4B · 84.9MiniCPM-V-4.5 · 8B · 82.6Qwen3.5-2B · 2B · 80.9TokenVL-2B · 2B · 78.1HunyuanOCR · 0.9B · 76.8HunyuanOCR1.5 · 0.9B · 76.5Qwen3.5-0.8B · 0.8B · 72.9Gemma-4-E4B-it · 4B · 61.4Gemma-4-E2B-it · 2B · 57.3Xiaomi-OCR-0 · 0.8B · 83.2Xiaomi-OCR-00.8B · 83.2
★ Xiaomi-OCR-0Colors & shapes identify models

Hover or select a point to view its reported scores.

02

Score comparison

Arithmetic mean of five benchmarks · higher is better

MODEL DETAILS
Xiaomi-OCR-0

0.8B parameters

83.2
DocVQA
93.1
InfoVQA
75.1
ChartQA
84.6
OCRBench
84.6
TextVQA
78.6
View detailed scores
Selected comparisons from the technical report. Column arrows indicate the preferred direction.
ModelSizeDocVQAInfoVQAChartQAOCRBenchTextVQAOverall (5-benchmark mean)
Xiaomi-OCR-00.8B93.175.184.684.678.683.2
Qwen3.5-4B4B94.480.482.486.680.884.9
MiniCPM-V-4.58B84.969.687.489.082.282.6
Qwen3.5-2B2B92.472.477.085.976.980.9
TokenVL-2B2B89.961.081.182.176.478.1
GPT-5.2-91.784.057.080.772.877.2
HunyuanOCR0.9B86.861.678.586.071.176.8
HunyuanOCR1.50.9B87.655.278.386.175.176.5
Qwen3.5-0.8B0.8B88.560.369.577.968.372.9
Gemma-4-E4B-it4B79.247.338.476.066.161.4
Gemma-4-E2B-it2B73.838.142.672.459.757.3
Gemini 3.0 Pro-----57.294.0----
Mini-Monkey2B87.460.176.5------
TextHawk27B89.667.881.4--75.1--

Our Real5, OmniDocBench v1.6, and Wild results use two-stage inference: PP-DocLayoutV3 detects regions, then Xiaomi-OCR-0 recognizes the crops and assembles the page. The 0.8B parameter count covers the VLM and excludes the external layout detector. Scores and model sizes are from the accompanying report; each tab uses its own benchmark metric.

03 / DATA & TRAINING

Data and training.

Our automated parsing data engine supports an approximately 170M-sample OCR-centric corpus. Training builds spatial alignment, broad task coverage, and task-specific optimization.

The parsing data engine

AUTOMATED ANNOTATION PIPELINE
Automated annotation pipeline: diverse document samples are laid out and detected, annotations are compared across heterogeneous experts, and difficult samples are refined through rendering and verification.
The pipeline combines broad document coverage, layout detection, expert consensus, and render-guided refinement.

Targeted synthesis for remaining gaps

COVERAGE-DRIVEN + FAILURE-DRIVEN
01 / COVERAGE-DRIVEN

Underrepresented patterns

Rare layouts, structures, writing styles, domains, and character combinations.

02 / FAILURE-DRIVEN

Verified recurring errors

Compare model predictions with trusted annotations, then diagnose repeated failure patterns.

Both routes select a concrete generation target
03 / GENERATE

Compose a recipe

Vary templates, text, visual style, backgrounds, and degradation to create image–annotation pairs.

04 / QUALITY CONTROL

Check each pair

Verify rendering and task rules, remove duplicates, and reject or regenerate invalid samples.

05 / RECIPE VALIDATION

Train and assess

Mix accepted samples into SFT; check frozen hard cases and broader regression results.

Review samples and evaluation results to accept the recipe, adjust it, or revisit the diagnosis.
Coverage targets expand what is scarce; failure targets test whether generated examples address verified model weaknesses.

THE TRAINING RECIPE

A foundation for
reading and understanding.

Starting from Qwen3.5-0.8B.

  1. 1
    ANCHOR

    Q-Mask text anchoring

    Align text content with its spatial location using token and mask supervision.

  2. 2
    EXPAND

    Continued pretraining

    Train on document parsing and OCR-centric understanding tasks in a shared autoregressive format.

  3. 3
    OPTIMIZE

    Mixed-task reinforcement learning

    Optimize with verifiable task rewards using a DAPO-style objective built on GRPO.

The report also compares Mix-RL with multi-teacher on-policy distillation (MOPD) as alternative post-training methods.

04 / RESEARCH INSIGHTS

Three findings from the experiments.

TAKEAWAY 01 / EMERGENT TRANSFER

Understanding benefits emerge with sufficient parsing training.

Absolute parsing scores at 9%, 29%, and 50% parsing-data stages
Three measured parsing-data stages; changes are relative to the baseline at each stage.

Understanding supervision hurts parsing at the early stage, but helps at the middle and late stages. Adding the same understanding pool changes Overall by -0.330 at the early stage and +0.639 at the middle stage. The late-stage ablation also shows positive transfer (+0.0749).

The stages use 9%, 29%, and 50% of all parsing data. Early training uses region-level data only. At the middle stage, page-level supervision adds +0.211 before understanding data is introduced, with about one page-level sample for every five region-level samples. The late stage keeps that ratio.

TAKEAWAY 02 / MODEL CAPACITY

Understanding places greater demands on model capacity.

The 4B teacher minus 0.8B Mix-RL gap is 0.15 parsing points and 4.9 understanding points.
4B task teachers versus the 0.8B Mix-RL model.

The teacher–student gap is 0.15 points on parsing and 4.9 points on understanding. Within each task, both models use the same amount of training data; each teacher specializes in one domain, while Mix-RL trains on both.

The larger understanding gap, together with emergent positive transfer, supports building a parsing foundation before placing greater emphasis on understanding.

TAKEAWAY 03 / POST-TRAINING COMPUTE

Mix-RL and MOPD training trajectories.

● Mix-RL● MOPD
OmniDocBench v1.6 Overall96.496.696.800.20.40.60.81Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.0238 · score 96.3866Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.1009 · score 96.4627Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.1290 · score 96.4913Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.1565 · score 96.5461Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.1873 · score 96.5000Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.2178 · score 96.5309Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.2492 · score 96.5636Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.2829 · score 96.5673Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.3193 · score 96.6668Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.3558 · score 96.7448Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.3932 · score 96.6440Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.4311 · score 96.5584Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.4705 · score 96.5961Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.5114 · score 96.5758Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.5536 · score 96.6914Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.5944 · score 96.6511Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.6348 · score 96.7487Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.6758 · score 96.7167Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.7222 · score 96.7700Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.7673 · score 96.8277Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.8134 · score 96.7442Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.8586 · score 96.6833Mix-RL · OmniDocBench v1.6 Overall · relative compute 0.9516 · score 96.7269Mix-RL · OmniDocBench v1.6 Overall · relative compute 1.0000 · score 96.7681MOPD · OmniDocBench v1.6 Overall · relative compute 0.0144 · score 96.3895MOPD · OmniDocBench v1.6 Overall · relative compute 0.0193 · score 96.5108MOPD · OmniDocBench v1.6 Overall · relative compute 0.0244 · score 96.4006MOPD · OmniDocBench v1.6 Overall · relative compute 0.0294 · score 96.4124MOPD · OmniDocBench v1.6 Overall · relative compute 0.0344 · score 96.4018MOPD · OmniDocBench v1.6 Overall · relative compute 0.0395 · score 96.4087MOPD · OmniDocBench v1.6 Overall · relative compute 0.0451 · score 96.4342MOPD · OmniDocBench v1.6 Overall · relative compute 0.0499 · score 96.4323MOPD · OmniDocBench v1.6 Overall · relative compute 0.0549 · score 96.4326MOPD · OmniDocBench v1.6 Overall · relative compute 0.0600 · score 96.4834MOPD · OmniDocBench v1.6 Overall · relative compute 0.0645 · score 96.5440MOPD · OmniDocBench v1.6 Overall · relative compute 0.0695 · score 96.5315MOPD · OmniDocBench v1.6 Overall · relative compute 0.0744 · score 96.5372MOPD · OmniDocBench v1.6 Overall · relative compute 0.0793 · score 96.4775MOPD · OmniDocBench v1.6 Overall · relative compute 0.0842 · score 96.6259MOPD · OmniDocBench v1.6 Overall · relative compute 0.0890 · score 96.6105MOPD · OmniDocBench v1.6 Overall · relative compute 0.0940 · score 96.5727MOPD · OmniDocBench v1.6 Overall · relative compute 0.0987 · score 96.6253MOPD · OmniDocBench v1.6 Overall · relative compute 0.1040 · score 96.5709MOPD · OmniDocBench v1.6 Overall · relative compute 0.1089 · score 96.6113MOPD · OmniDocBench v1.6 Overall · relative compute 0.1135 · score 96.5975MOPD · OmniDocBench v1.6 Overall · relative compute 0.1182 · score 96.6693MOPD · OmniDocBench v1.6 Overall · relative compute 0.1230 · score 96.6927MOPD · OmniDocBench v1.6 Overall · relative compute 0.1279 · score 96.7417MOPD · OmniDocBench v1.6 Overall · relative compute 0.1326 · score 96.6778MOPD · OmniDocBench v1.6 Overall · relative compute 0.1376 · score 96.6639MOPD · OmniDocBench v1.6 Overall · relative compute 0.1424 · score 96.6889MOPD · OmniDocBench v1.6 Overall · relative compute 0.1472 · score 96.6330MOPD · OmniDocBench v1.6 Overall · relative compute 0.1521 · score 96.6208MOPD · OmniDocBench v1.6 Overall · relative compute 0.1569 · score 96.6335MOPD · OmniDocBench v1.6 Overall · relative compute 0.1616 · score 96.7146MOPD · OmniDocBench v1.6 Overall · relative compute 0.1666 · score 96.6665MOPD · OmniDocBench v1.6 Overall · relative compute 0.1713 · score 96.7092MOPD · OmniDocBench v1.6 Overall · relative compute 0.1761 · score 96.6565MOPD · OmniDocBench v1.6 Overall · relative compute 0.1809 · score 96.6115MOPD · OmniDocBench v1.6 Overall · relative compute 0.1860 · score 96.6398MOPD · OmniDocBench v1.6 Overall · relative compute 0.1908 · score 96.6852MOPD · OmniDocBench v1.6 Overall · relative compute 0.1955 · score 96.6350MOPD · OmniDocBench v1.6 Overall · relative compute 0.2004 · score 96.6517MOPD · OmniDocBench v1.6 Overall · relative compute 0.2048 · score 96.6701Relative cumulative training compute

Hover or select a checkpoint to view its score and relative compute.

OCR-VQA · five-benchmark mean81828300.20.40.60.81Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.0238 · score 80.8980Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.0482 · score 81.5020Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.0740 · score 82.1800Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.1009 · score 82.4500Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.1290 · score 82.8000Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.1565 · score 82.6660Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.1873 · score 83.0800Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.2178 · score 82.9480Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.2492 · score 83.1840Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.2829 · score 83.0060Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.3193 · score 82.8480Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.3558 · score 82.8720Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.3932 · score 82.9640Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.4311 · score 83.1720Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.4705 · score 82.7860Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.5114 · score 83.0940Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.5536 · score 82.9980Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.5944 · score 83.2640Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.6348 · score 83.0740Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.6758 · score 83.2080Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.7222 · score 83.0720Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.7673 · score 83.2140Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.8134 · score 83.2240Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.8586 · score 83.0940Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.9038 · score 83.2200Mix-RL · OCR-VQA · five-benchmark mean · relative compute 0.9516 · score 82.9320MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0054 · score 81.0720MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0099 · score 80.9920MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0144 · score 81.1420MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0193 · score 80.9900MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0244 · score 80.9200MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0294 · score 81.3260MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0344 · score 81.3100MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0395 · score 81.4340MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0451 · score 81.3880MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0499 · score 81.4280MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0549 · score 81.8780MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0600 · score 81.7420MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0645 · score 81.4740MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0695 · score 81.6460MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0744 · score 81.8360MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0793 · score 81.8660MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0842 · score 82.3480MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0890 · score 82.0380MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0940 · score 81.7640MOPD · OCR-VQA · five-benchmark mean · relative compute 0.0987 · score 82.0000MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1040 · score 82.3640MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1089 · score 82.2740MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1135 · score 82.0920MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1182 · score 82.2420MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1230 · score 82.1580MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1279 · score 81.9540MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1326 · score 82.2960MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1376 · score 82.1320MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1424 · score 82.3320MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1472 · score 82.3060MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1521 · score 82.2300MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1569 · score 82.1680MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1666 · score 82.5400MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1713 · score 82.4380MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1761 · score 82.4580MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1809 · score 82.6600MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1860 · score 82.4420MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1908 · score 82.5580MOPD · OCR-VQA · five-benchmark mean · relative compute 0.1955 · score 82.7080MOPD · OCR-VQA · five-benchmark mean · relative compute 0.2004 · score 82.6520MOPD · OCR-VQA · five-benchmark mean · relative compute 0.2048 · score 82.5580Relative cumulative training compute

Hover or select a checkpoint to view its score and relative compute.

Replay the training progression; hover or select any checkpoint for exact values. The horizontal axis is normalized training compute.

The two runs use different total compute budgets; the curves show the observed scores during training. MOPD training was stopped after monitored scores began to fluctuate.

Observed runs: Mix-RL peaks at 96.8277 on parsing and 83.264 on OCR-VQA; MOPD peaks at 96.7417 and 82.708, respectively. OCR-VQA uses the fixed five-benchmark mean.

05 / AGENT SKILL

Use OCR in your agent.

AGENT SKILL

Install from your agent.

Send this one-line instruction to your coding agent:

PASTE INTO YOUR AGENT
Read and execute https://raw.githubusercontent.com/SeerRay-Lab/Xiaomi-OCR-0/main/SKILL.md
⚠️ Notice:

This Skill may lead your agent to download model weights, set up an SGLang or vLLM environment, and configure an MCP server. If you do not agree to this setup, do not run the command. Read the setup details on GitHub first.

SETUP DETAILSGitHub repository ↗skills/xiaomi-ocr/SKILL.md ↗

06 / INFERENCE

Inference and evaluation.

Choose an inference path, preserve the evaluation settings, and keep the predictions needed to check each score.

INFERENCE MODES

One model, two parsing paths.

END TO END

A full page in.
Markdown out.

The VLM reads the complete image and returns text, OTSL tables, and LaTeX formulas in reading order.

Page image→Xiaomi-OCR-0→Markdown
pipeline/e2e_img2md.py
TWO STAGE

Locate regions.
Recognize and assemble.

PP-DocLayoutV3 detects regions. The VLM recognizes each crop with a task prompt, then the pipeline assembles the page.

Layout→Region OCR→Assembly
pipeline/twostage_img2md.py
Read scores with their inference mode.

The headline parsing results use two-stage inference. The 0.8B size describes Xiaomi-OCR-0; it excludes the layout detector. The report also evaluates direct full-page inference.

Report inference-mode comparison
BenchmarkInferenceOverall ↑Formula CDM ↑Table TEDS ↑
OmniDocBench v1.6End to end95.192696.455392.5325
OmniDocBench v1.6Two stage96.827798.504495.1087
Wild-OmniDocBenchEnd to end86.480690.073480.6184
Wild-OmniDocBenchTwo stage87.943889.548586.6129

BEFORE YOU START

Use the matching checkpoint and environment.

If the Xiaomi-OCR Skill has already prepared the code, runtime, and model on your machine, you can skip this setup. Otherwise, start with the GitHub code and follow the installation guide: clone the code, choose and install SGLang or vLLM, then download the model checkpoint from Hugging Face. The inference examples below assume the local model service is already running.

  • Use the repository setup instructions to install the code dependencies and your chosen inference backend; start an OpenAI-compatible endpoint at 127.0.0.1:8000 with model name SeerRay-Lab/Xiaomi-OCR-0.
  • Download the Xiaomi-OCR-0 checkpoint from Hugging Face and configure your runtime to load its local model directory.
  • For two-stage region parsing, also install PaddleX and download the PP-DocLayoutV3 layout weights as described in the repository guide. Full-page inference does not require this layout model.

Create images.jsonl with one image per line. Use absolute paths and unique filename stems, since output files are named after each input image.

images.jsonl
{"images": ["/absolute/path/to/page-001.png"]}
{"images": ["/absolute/path/to/page-002.png"]}

RUN INFERENCE

Choose the command for your setup.

Install requirements.txt from the repository root, then use a fresh output directory for each configuration. Module invocation keeps the repository’s shared postprocessing imports available.

Full-page inference

From the repository root
python -m pipeline.e2e_img2md \
  --input-jsonl images.jsonl \
  --output_dir outputs/e2e \
  --served_model_name SeerRay-Lab/Xiaomi-OCR-0 \
  --port 8000 \
  --task document \
  --max_tokens 16384

Layout detection and region recognition

Local PP-DocLayoutV3 + existing OCR endpoint
python -m pipeline.layout_detect_paddlex \
  --input_jsonl images.jsonl \
  --output_jsonl layout.jsonl \
  --gpu_ids 0

python -m pipeline.twostage_img2md \
  --input-jsonl images.jsonl \
  --output_dir outputs/two-stage \
  --layout-jsonl layout.jsonl \
  --served_model_name SeerRay-Lab/Xiaomi-OCR-0 \
  --port 8000 \
  --temperature 0.0 \
  --max_tokens 16384

If you already exported a PaddleX layout, skip the first command and pass its path with --layout-jsonl. Install a compatible PaddlePaddle wheel and requirements-paddlex.txt before layout detection; see the GitHub installation guide.

Prompts, OTSL, and repetition handling

The full-page and table-region prompts request OTSL for tables. OTSL is the model output format. The shared postprocessor converts tables to HTML inside the returned Markdown; MCP and the Demo apply it automatically. If repetitive-output retry is enabled, record that setting with the inference configuration and preserve the first-pass output for comparison.

Thinking is off unless --think is supplied. Context capacity must cover the image, prompt, and generated output; --max_tokens alone does not set server context capacity.

OUTPUTS

Keep the evidence behind the score.

OutputWhat to keep it for
*.mdOne parsed document per input image, used for qualitative checks and benchmark evaluation.
results.jsonlTwo-stage layout regions, recognition outputs, and assembled Markdown.
primary_results.jsonlTwo-stage first-pass results when repetition retries are enabled.
repetition_retry_report.jsonRetry diagnostics; the full-page script always writes this report, while the two-stage script writes it when retries are enabled.

postprocess/assemble_markdown.py can rebuild pages from saved region results. Preserve the assembly settings: text, formula, image, and reading-order handling can affect the evaluated output.

EVALUATION

Compare like with like.

BenchmarkProtocol to preserve
OmniDocBench v1.6Official v1.6 evaluator and split; report Overall, Text Edit, Formula CDM, Table TEDS, TEDS-S, and Reading Order Edit. Record whether inference is end to end or two stage.
Real5-OmniDocBenchRetain the five acquisition conditions: scanning, warping, screen photo, illumination, and skew.
Wild-OmniDocBenchUse the physical-scene document benchmark and its matching evaluator, with the same parsing mode as the target score.
OCR-centric understandingOverall is the mean of DocVQA, InfoVQA, ChartQA, OCRBench, and TextVQA on a 0–100 scale. Report Overall only when all five scores are available.

The inspected repository provides parsing pipelines and the Agent runtime. It does not yet include a complete, pinned benchmark suite or the OCR-VQA evaluation commands. This guide documents the supported inference interface; exact score reproduction still needs those release artifacts.

Explore the reported results ↗

TASK INSTRUCTIONS

Official prompts.

Document · full-page parsing
Default pipeline prompt
Extract all information from the main body of the document image and represent it in markdown format, ignoring headers and footers. Tables should be expressed in OTSL format, formulas in the document should be represented using LATEX format, and the parsing should be organized according to the reading order.
Text, formulas, and tables
Text
Extract the text in the image.
Formula
Identify the formula in the image and represent it using LATEX format.
Table
Parse the table in the image into OTSL.
Key information extraction · KIE

Provide the fields you want as a JSON schema. The prompt starts with this instruction:

KIE prompt template
Extract key information in the image

Please output the key information in JSON format according to the following schema:
{
    "date": "",
    "seller_address": "",
    "seller_gst_id": "",
    "seller_name": "",
    "total_amount": "",
    "total_tax": ""
}

Replace the example keys with the fields required for your document. When using MCP, pass field names to kie_image; the service builds the prompt and cleans the JSON response.

TECHNICAL REPORT

Xiaomi-OCR-0 Technical Report

arXiv:2609.36136 ↗

Copy the BibTeX entry to cite this work.

BIBTEX
@misc{chen2026xiaomiocr0technicalreport,
  title={Xiaomi-OCR-0 Technical Report},
  author={Xin Chen and Anan Du and Feng Feng and Pei Fu and Jian Luan and Longwei Xu and Shaojie Zhang and Hang Li and Heng Qu and Cheng Tan},
  year={2026},
  eprint={2609.36136},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.36136},
}