THE TECHNIQUE OF EXTRACTION TEXT AREAS ON SCANNED DOCUMENT IMAGE USING LINEAR FILTRATION
Journal: Applied Aspects of Information Technology (Vol.02, No. 03)Publication Date: 2019-09-03
Authors : Ishchenko Alesya; Nesteryuk Alexandr; Polyakova Marina;
Page : 206-215
Keywords : image segmentation; text areas; scanned document; linear filtering; image processing;
Abstract
The method of selection of text areas on the image of the scanned document from the background is proposed. Text areas of the image have approximately the same intensity values inside these areas. Therefore, linear filtering and threshold image transformation are used. Linear filtering allows you to smooth out the intensity values of pixels inside homogeneous areas. In the case of a threshold transformation, the threshold value is used, which makes it possible to isolate homogeneous areas of the im-age that make up the text fragments from the background.A study was conducted on the selection of a threshold value for highlight-ing homogeneous areas of text, which showed that the threshold value is better to choose among the pixel intensities at the base of the histogram peak, which corresponds to the background. It is proposed to select the threshold by the value of the second derivative for the image histogram after linear filtering. Therefore, the intensity of the local maximum of the histogram, which is closer than the other local maxima to the right end of the image intensity interval, is chosen as the threshold. For this purpose, an analysis of the histogram of the distribution of image pixel intensity values is carried out after linear filtering by rows and columns at each step. Testing of the proposed method of separating textual image areas was carried out for segmentation of textual images of scanned archival newspapers from the MediaTeam documents database at the University of Oulu (Finland).The proposed method of extract-ing text fragments from the background using linear filtering and threshold conversion allowed to improve the quality of selection of these areas compared to the similar method in the percentage of correct recognition of text areas by 12 %, which is important for the task of image segmentation.
Other Latest Articles
- ANALYSIS AND SYNTHESIS OF THE RESULTS OF COMPLEX EXPERIMENTAL RESEARCH ON REENGINEERING OF OPEN CAD SYSTEMS
- MODELS AND METHODS OF INTELLECTUAL ANALYSIS FOR MEDICAL-SOCIOLOGICAL MONITORING’S DATA BASED ON THE NEURAL NETWORK WITH A COMPETITIVE LAYER
- PROOF-OF-GREED APPROACH IN THE NXT CONSENSUS
- COMPOSITIONAL METHOD OF FPGA PROGRAM CODE INTEGRITY MONITORING BASED ON THE USAGE OF DIGITAL WATERMARKS
- MODELS BASED ON CONFORMAL PREDICTORS FOR DIAGNOSTIC SYSTEMS IN MEDICINE
Last modified: 2019-09-03 01:52:54