- AutorIn
- M. Sc. Jannes Nitzsche Universität Leipzig#Deutsche Telekom MMS
- Titel
- Multimodal Approaches for Information Extraction from Industrial Product Labels
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:15-qucosa2-1048221
- Erstveröffentlichung
- 2026
- Datum der Einreichung
- 30.03.2026
- Datum der Verteidigung
- 18.02.2026
- Abstract (EN)
- The automatic extraction of structured information from industrial product labels is a critical task in logistics and warehouse automation, yet it remains challenging due to high label variability and degraded image conditions encountered in real-world environments. This thesis investigates and compares two fundamentally different extraction paradigms: Vision-Language Model (VLM) based end-to-end extraction using Qwen2-VL-7B, Gemma-3-12B and Mistral-Small-3.2-24B, and a two-stage Optical Character Recognition (OCR) based pipeline using DocTR and PaddleOCR combined with a custom contextualization algorithm. To support the evaluation, a synthetic dataset with automated annotations was created, extended with augmented variants simulating real-world degradations, and complemented by a realistic subset obtained by printing and photographing physical labels. A dedicated evaluation framework was developed to assess and compare all methods in terms of extraction quality and processing speed. The results show that Mistral-Small-3.2-24B achieves the highest overall extraction quality with an object F1-Score of 0.906 on the synthetic dataset and the greatest robustness to image degradation across all evaluated methods. PaddleOCR with English-only recognition offers competitive accuracy while being considerably faster than any VLM-based method. VLM-based methods are recommended when extraction quality, robustness to varying input conditions, or versatility across different label formats and use cases are the primary objectives. OCR-based extraction is better suited for real-time or resource-constrained deployments with stable input conditions.
- Freie Schlagwörter (EN)
- Vision-Language Model, Optical Character Recognition, Document Understanding, Logistics Automation, Information Extraction
- Klassifikation (DDC)
- 004
- BetreuerIn Hochschule / Universität
- Dr. Thomas Burghardt
- BetreuerIn - externe Einrichtung
- M. Sc. Ilja Dontsov
- Den akademischen Grad verleihende / prüfende Institution
- Universität Leipzig, Leipzig
- Sonstige beteiligte Institution
- ScaDS.AI, Leipzig
- Deutsche Telekom MMS, Dresden
- Version / Begutachtungsstatus
- publizierte Version / Verlagsversion
- URN Qucosa
- urn:nbn:de:bsz:15-qucosa2-1048221
- Veröffentlichungsdatum Qucosa
- 26.05.2026
- Dokumenttyp
- Masterarbeit / Staatsexamensarbeit
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis
CC BY 4.0- Inhaltsverzeichnis
I. Foundations 1. Introduction 1.1. Motivation 1.2. Objectives 1.3. Project Context 1.4. Structure of the Thesis 2. Fundamentals 2.1. Model Architectures 2.1.1. Convolutional Neural Networks 2.1.2. Transformer Models 2.2. Optical Character Recognition 2.3. Vision Language Models and Multimodal Transformers 2.3.1. Definition and General Concept 2.3.2. Architectural Variants 2.3.3. Image-Text Alignment and Multimodal Embeddings 2.3.4. Prompting Principles 2.3.5. Model Compression and Quantization 2.4. Document Understanding 2.4.1. Multimodal Document Representation 2.4.2. Document Visual Question Answering 2.5. Synthetic Data 2.6. Industrial Label Characteristics 2.7. Evaluation Metrics for Information Extraction 2.7.1. Character- and Word-Level Metrics 2.7.2. Information Group Extraction Metrics 2.7.3. Accuracy-Speed Trade-off and Multi-Dimensional Evaluation 3. Related Work 3.1. OCR-based Pipelines for Label and Document Understanding 3.2. Vision-Language Models and End-to-End Structured Extraction 3.3. Synthetic Dataset Generation and Image Augmentation 3.4. Robustness of Recognition Systems to Image Degradation 3.5. Research Gap II. Methodology 4. Dataset 4.1. Synthetic Label Generation 4.1.1. Generated Label Content 4.1.2. Layout Variability and Visual Structure 4.1.3. Annotation Output 4.2. Synthetic Data Augmentation 4.2.1. Motivation for Augmentation 4.2.2. Augmentation Framework: Augraphy 4.2.3. Applied Augmentations 4.3. Realistic Label Generation 4.3.1. Label Generation and Printing 4.3.2. Image Capture Setup 4.3.3. Image Processing and Preparation 4.3.4. Annotation Process 4.4. Dataset Statistics 4.4.1. Object Type Distribution 4.4.2. Label Complexity 4.4.3. Text Field Categories 4.4.4. Spatial Distribution of Annotations 5. Text Extraction 5.1. VLM-Based Text Extraction 5.1.1. Model Selection 5.1.2. Inference 5.1.3. Prompt Design 5.1.4. Hallucination Mitigation 5.2. OCR-Based Text Extraction 5.2.1. OCR Framework Selection 5.2.2. OCR Output Characteristics 5.2.3. Contextualization 6. Evaluation Strategy 6.1. OCR-Specific Evaluation 6.1.1. Text Aggregation 6.1.2. Unordered Word Error Rate 6.2. End-to-End Evaluation III. Results & Discussion 7. Evaluation Results 7.1. OCR-Specific Results 7.2. End-to-End Results 7.3. Common Errors and Failure Cases 8. Summary and Outlook 8.1. Summary 8.2. Limitations 8.3. Future Work