Document Reader System |
Author(s): |
| Kamble Akash Goroba , SVPMS COE MALEGAON BK; Atole Swapnali D., SVPMS COE MALEGAON BK; Mote Kajal B., SVPMS COE MALEGAON BK; Sudrik Rahul D., SVPMS COE MALEGAON BK |
Keywords: |
| Tresaract, Optical Character Recognition, Feature Extraction (OCR), Text To Speech (TTS), Feature Matching, Text Extraction, Character Extraction |
Abstract |
|
In today’s post, we will learn how to recognize text in images using an open source tool called Tesseract and OpenCV. The method of extracting text from images is also called Optical Character Recognition (OCR) or sometimes simply text recognition. In this paper an assistive system has been proposed which is useful for visually impaired or also normal person. It is the system which reads texual information present on papers and produce corresponding voice using OCR(Optical Character Recognition)and TTS(Text-To-Speech) system. Optical Character Recognition (OCR) is a system that provides a full alphanumeric recognition of printed or handwritten characters by simply scanning the text image. OCR system interprets the printed or handwritten characters image and converts it into corresponding editable text document. The text image is divided into regions by isolating each line, then individual characters with spaces. After character extraction, the texture and topological features like corner points, features of different regions, ratio of character area and convex area of all characters of text image are calculated. Previously features of each uppercase and lowercase letter, digit, and symbols are stored as a template. Based on the texture and topological features, the system recognizes the exact character using feature matching between the extracted character and the template of all characters as a measure of similarity. |
Other Details |
|
Paper ID: IJSRDV7I30786 Published in: Volume : 7, Issue : 3 Publication Date: 01/06/2019 Page(s): 1516-1518 |
Article Preview |
|
|
|
|
