Document Type

Article

Publication Date

1-2025

Publisher

Institute of Electrical and Electronics Engineers (IEEE)

Source Publication

IEEE Access

Source ISSN

2169-3536

Abstract

Information extraction from financial document images is crucial in computer vision and NLP, as financial data often exists in image or PDF format, enabling organizations to analyze and make informed business decisions using OCR advancements. The table contents of financial document images are one of the prominent structures to confine important portions of data of the document and many Deep learning-based methods have been proposed to detect Table regions inside document images. The shortcomings of the current approach are that it is bounded within the detection of the table region and struggles in cases such as handling different layouts and preserving the relation among the different attributes of the table. Therefore, in this work, we proposed an end-to-end architecture to extract information from Financial table images while preserving the column row structures of the attributes within the table. We divided the task into four modules and generated synthesized data with different augmentation techniques to overcome data scarcity challenges and boost the performance of the pipeline modules. In terms of information extraction, the proposed method acquired 85% accuracy in the target invoice dataset.

Comments

Published version. IEEE Access, Vol. 13 (2025): 17706-17723. DOI. © 2024 The Authors. Used with permission.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.

deshpande_17523acc.docx (1151 kB)
ADA Accessible Version

Share

COinS