Enhancing Scanned Documents Using Multi-Stage Modular Processing
Pages
40-52Abstract
Images of scanned documents are often vulnerable to complicated degradations, including uneven illumination, shadows, and haze, greatly reducing text readability and the performance of Optical Character Recognition (OCR) systems. This paper presents a multi-step modular model that aims at improving the quality of documents by working on each of these distortions individually. The system architecture comprises two specific modules that include the shadow removal module and the haze/blur reduction pipeline. The shadow removal algorithm employs background estimation with median filtering and illumination normalization, and then a refinement step with the DRUNet pre-trained model to avoid residual noise, at the expense of fine textual features. In addition to that, the haze reduction system utilizes a soft-enhancement technique employing Contrast Limited Adaptive Histogram Equalization (CLAHE) combined with Gaussian smoothing to achieve local contrast enhancement with no visual artefacts. Experimental results indicate that the proposed modular approach produces a neater background and sharper edge definition in the text than the traditional single-stage method. MSE, PSNR, SSIM, and GMSD measurements are quantitative analyses of the framework's effectiveness, and the results indicate that all document types have similar improvements.
Keywords:
References
- [1] B. Wang, Z. Wang, W. Liu, X. Huang, C. L. P. Chen, and Y. Zhao, “DDSR-net: Direct document shadow removal leveraging multi-scale attention,” Mach. Intell. Res., vol. 22, no. 3, pp. 452–465, 2025, doi: https://doi.org/10.1007/s11633-024-1522-4.
- [2] W. Liu, B. Wang, J. Zheng, and W. Wang, “Shadow removal of text document images using background estimation and adaptive text enhancement,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. doi: DOI: 10.1109/ICASSP49357.2023.10096115.
- [3] M. Guijarro, J. Bayon, D. Mart’in-Carabias, and J. Recas, “A Multi-Stage Method for Logo Detection in Scanned Official Documents Based on Image Processing,” Algorithms, vol. 17, no. 4, p. 170, 2024, doi: https://doi.org/10.3390/a17040170.
- [4] P. Dubey and D. R. Shashikumar, “A Multi-Stage Deep Learning Model for the Enhancement of the Quality of Camera-Captured Document Images,” Eng. Technol. & Appl. Sci. Res., vol. 15, no. 6, pp. 29964–29970, 2025, doi: https://doi.org/10.48084/etasr.14469.
- [5] D. Gao, W. Liu, S. Chen, J. Qiu, X. Mei, and B. Wang, “Document Image Shadow Removal Based on Illumination Correction Method,” Algorithms, vol. 18, no. 8, p. 468, 2025, doi: https://doi.org/10.3390/a18080468.
- [6] Y. Jin, Y. Wang, Q. Zhong, K. C. Jin-Chun, K. Z. Ke, and D. MacDonald, “Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models,” arXiv Prepr. arXiv2510.23066, 2025, doi: https://doi.org/10.48550/arXiv.2510.23066.
- [7] Z. Anvari and V. Athitsos, “A survey on deep learning based document image enhancement,” arXiv Prepr. arXiv2112.02719, 2021, doi: https://doi.org/10.48550/arXiv.2112.02719.
- [8] M. Karayaka, U. Muhammad, J. Laaksonen, M. Z. Hoque, and T. Seppänen, “A Dual-Domain Convolutional Network for Hyperspectral Single-Image Super-Resolution,” arXiv Prepr. arXiv2512.09546, 2025, doi: https://doi.org/10.48550/arXiv.2512.09546.
- [9] B. Wang, C. Li, W. Zou, Y. Zhang, X. Chen, and C. L. P. Chen, “A comprehensive survey on shadow removal from document images: datasets, methods, and opportunities,” Vicinagearth, vol. 2, no. 1, p. 1, 2025, doi: https://doi.org/10.1007/s44336-024-00010-9.
- [10] Y. Cui, Y. Tao, L. Jing, and A. Knoll, “Strip attention for image restoration,” in International Joint Conference on Artificial Intelligence, IJCAI, 2023. doi: https://doi.org/10.24963/ijcai.2023/72.
- [11] C. Wang et al., “Promptrestorer: A prompting image restoration method with degradation perception,” Adv. Neural Inf. Process. Syst., vol. 36, pp. 8898–8912, 2023, doi: https://dl.acm.org/doi/abs/10.5555/3666122.3666512.
- [12] Z. Chen, Z. He, and Z.-M. Lu, “DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attention,” IEEE Trans. image Process., vol. 33, pp. 1002–1015, 2024, doi: DOI: 10.1109/TIP.2024.3354108.
- [13] X. Chen et al., “Unpaired deep image dehazing using contrastive disentanglement learning,” in European conference on computer vision, 2022, pp. 632–648. doi: DOI https://doi.org/10.1007/978-3-031-19790-1_38.
- [14] G. L. N. Vanguri, S. K. Swain, and M. V. Krishna, “An Investigative Study on Deep Learning-Based Image Dehazing Techniques,” in International Conference on Computational Innovations and Emerging Trends (ICCIET-2024), 2024, pp. 87–97. doi: DOI 10.2991/978-94-6463-471-6_9.
- [15] Y. Zhou, S. Zuo, Z. Yang, J. He, J. Shi, and R. Zhang, “A review of document image enhancement based on document degradation problem,” Appl. Sci., vol. 13, no. 13, p. 7855, 2023, doi: https://doi.org/10.3390/app13137855.
- [16] Z. Zhou et al., “Docdeshadower: Frequency-aware transformer for document shadow removal,” in 2024 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2024, pp. 2468–2473. doi: DOI: 10.1109/SMC54092.2024.10831480.
- [17] B. Wang and C. L. P. Chen, “Local water-filling algorithm for shadow detection and removal of document images,” Sensors, vol. 20, no. 23, p. 6929, 2020, doi: https://doi.org/10.3390/s20236929.
- [18] X. Chen, X. Cun, C.-M. Pun, and S. Wang, “Shadocnet: Learning spatial-aware tokens in transformer for document shadow removal,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. doi: DOI: 10.1109/ICASSP49357.2023.10095403.
- [19] A. Fathallah, M. El Yacoubi, and N. Ben Amara, “EHDI: enhancement of historical document images via generative adversarial network,” in 18th International Conference on Computer Vision Theory and Applications (VISAPP), 2023, pp. 238–245. doi: DOI : 10.5220/0011662700003417.
- [20] H. Neji, T. M. Hamdani, M. B. Halima, J. Nogueras-Iso, and A. M. Alimi, “Blur2sharp: A gan-based model for document image deblurring,” 2021. doi: DOI: 10.2991/IJCIS.D.210407.001.
- [21] Z. Yang et al., “Docdiff: Document enhancement via residual diffusion models,” in Proceedings of the 31st ACM international conference on multimedia, 2023, pp. 2795–2806. doi: https://doi.org/10.1145/3581783.3611730.
- [22] A. A. Abdulmajeed and N. N. Saleem, “Pre-processing and enhancement techniques for COVID-19 X-ray and CT-scan medical images,” in AIP Conference Proceedings, 2022, p. 50011. doi: https://doi.org/10.1063/5.0121956.
- [23] N. T. Saeed and H. M. Ahmed, “Building a Real-Time System to Monitor Students Electronically Based on Digital Images of Face Movement,” ICOASE 2022 - 4th Int. Conf. Adv. Sci. Eng., pp. 83–88, 2022, doi: 10.1109/ICOASE56293.2022.10075587.
- [24] H. S. Abdulla, A. S. Shaheen, and N. M. Isaac, “Effectiveness of Image Curvelet Transform Coefficients for Image Denoising”, doi: https://doi.org/10.33899/csmj.2024.146534.1105.
- [25] J. Xu et al., “General information metrics for improving AI model training efficiency,” Artif. Intell. Rev., vol. 58, no. 9, p. 289, 2025, doi: https://doi.org/10.1007/s10462-025-11281-z.
- [26] M. D. Badiuzzaman, “Unpacking the metrics: a critical analysis of the 2025 QS World University Rankings using Australian university data,” in Frontiers in Education, 2025, p. 1619897. doi: https://doi.org/10.3389/feduc.2025.1619897.
- [27] K. Zhang, Y. Li, W. Zuo, L. Zhang, L. Van Gool, and R. Timofte, “Plug-and-play image restoration with deep denoiser prior,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 10, pp. 6360–6376, 2021, doi: DOI: 10.1109/TPAMI.2021.3088914.
- [28] R. Vimal, A. Kumar, and A. Tiwari, “A novel compression method for Vectorcardiogram signal using hybrid autoencoder model,” Biomed. Signal Process. Control, vol. 113, p. 108796, 2026, doi: https://doi.org/10.1016/j.bspc.2025.108796.
- [29] A. Dziembowski, W. Nowak, and J. Stankowski, “IV-SSIM-The structural similarity metric for immersive video,” Appl. Sci., vol. 14, no. 16, p. 7090, 2024, doi: https://doi.org/10.3390/app14167090.
- [30] A. Di Marino, V. Bevilacqua, E. Di Nardo, A. Ciaramella, I. De Falco, and G. Sannino, “A new Image Similarity Metric for a Perceptual and Transparent Geometric and Chromatic Assessment,” arXiv Prepr. arXiv2601.19680, 2026, doi: https://doi.org/10.48550/arXiv.2601.19680.
Identifiers
Download this PDF file
Statistics
How to Cite
Copyright and Licensing

This work is licensed under a Creative Commons Attribution 4.0 International License.





