Skip to main navigation Skip to search Skip to main content

Modular OCR Using Web Scraping Data

  • Guy Gisfan
  • , Eli O. David
  • , Nathan S. Netanyahu
  • Bar-Ilan University
  • College of Law and Business

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper explores the effectiveness of traditional (modular) Optical Character Recognition (OCR) pipelines compared to Large Vision Language Models (LVLMs), particularly in terms of robustness, accuracy, and resource efficiency. We propose leveraging web-generated datasets to train OCR systems, highlighting the rich diversity of layouts, styles, and linguistic variations offered by web content. Our approach shows that web scraping with smart augmentations can generate diverse OCR training datasets for training modular OCR. Acquiring website screenshots for modular OCR training has yet to be explored and requires precise word localization for training the word detection model. This differs from many LVLMs, which are mostly trained end-to-end to extract full text. Experimental evaluations demonstrate that even though we trained our OCR pipeline on design website templates for developers rather than real public websites, we achieved competitive and even superior results with the state-of-the-art LVLM-based models and superior results in noisy and distorted scenarios, while requiring fewer computational resources for training and inference. These findings underscore the long-lasting relevance of modular OCR systems in diverse and resource-constrained settings. Although LVLMs present advantages in handling diverse and generalized tasks, their use is unnecessary for the OCR task, where modular pipelines excel in terms of efficiency and performance.

Original languageEnglish
Title of host publicationDocument Analysis and Recognition – ICDAR 2025 Workshops - Proceedings
EditorsLianwen Jin, Richard Zanibbi, Veronique Eglin
PublisherSpringer Science and Business Media Deutschland GmbH
Pages194-210
Number of pages17
ISBN (Print)9783032093707
DOIs
StatePublished - 2026
EventInternational Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025 - Wuhan, China
Duration: 20 Sep 202521 Sep 2025

Publication series

NameLecture Notes in Computer Science
Volume16225 16226 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

ConferenceInternational Workshops co-located with the 19th International Conference on Document Analysis and Recognition, ICDAR 2025
Country/TerritoryChina
CityWuhan
Period20/09/2521/09/25

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 8 - Decent Work and Economic Growth
    SDG 8 Decent Work and Economic Growth
  2. SDG 12 - Responsible Consumption and Production
    SDG 12 Responsible Consumption and Production

Keywords

  • Large Vision Language Model (LVLM)
  • Optical Character Recognition (OCR)
  • Web-based data generation
  • Word detection
  • Word recognition

Fingerprint

Dive into the research topics of 'Modular OCR Using Web Scraping Data'. Together they form a unique fingerprint.

Cite this