Abstract
This paper explores the effectiveness of traditional (modular) Optical Character Recognition (OCR) pipelines compared to Large Vision Language Models (LVLMs), particularly in terms of robustness, accuracy, and resource efficiency. We propose leveraging web-generated datasets to train OCR systems, highlighting the rich diversity of layouts, styles, and linguistic variations offered by web content. Our approach shows that web scraping with smart augmentations can generate diverse OCR training datasets for training modular OCR. Acquiring website screenshots for modular OCR training has yet to be explored and requires precise word localization for training the word detection model. This differs from many LVLMs, which are mostly trained end-to-end to extract full text. Experimental evaluations demonstrate that even though we trained our OCR pipeline on design website templates for developers rather than real public websites, we achieved competitive and even superior results with the state-of-the-art LVLM-based models and superior results in noisy and distorted scenarios, while requiring fewer computational resources for training and inference. These findings underscore the long-lasting relevance of modular OCR systems in diverse and resource-constrained settings. Although LVLMs present advantages in handling diverse and generalized tasks, their use is unnecessary for the OCR task, where modular pipelines excel in terms of efficiency and performance.
| Original language | English |
|---|---|
| Title of host publication | Document Analysis and Recognition – ICDAR 2025 Workshops - Proceedings |
| Editors | Lianwen Jin, Richard Zanibbi, Veronique Eglin |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 194-210 |
| Number of pages | 17 |
| ISBN (Print) | 9783032093707 |
| DOIs | |
| State | Published - 2026 |
| Event | International Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025 - Wuhan, China Duration: 20 Sep 2025 → 21 Sep 2025 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Volume | 16225 16226 LNCS |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | International Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025 |
|---|---|
| Country/Territory | China |
| City | Wuhan |
| Period | 20/09/25 → 21/09/25 |
Bibliographical note
Publisher Copyright:© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 8 Decent Work and Economic Growth
-
SDG 12 Responsible Consumption and Production
Keywords
- Large Vision Language Model (LVLM)
- Optical Character Recognition (OCR)
- Web-based data generation
- Word detection
- Word recognition
Fingerprint
Dive into the research topics of 'Modular OCR Using Web Scraping Data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver