A Hybrid Model to Automatically Extract Text From Images |
Author(s): |
| H.Thaheera MCA., MPhil , Jamal Mohammed College; Professor A.Abdul Samathu M.Sc.,M.Phil.,PGDCA.,M.Tech, Jamal Mohamed College,Trichy |
Keywords: |
| Text Mining, Text Extraction from Natural Scene Images, text detection, text localization - CRF |
Abstract |
|
Text mining also referred to as text data mining, which is equivalent to text analytics, refers to the process of deriving high-quality information from text and such information is typically derived through the devising of patterns and trends through means such as statistical pattern learning. This Text mining usually involves the process of structuring the input text like parsing, along with the addition of some derived language based features and the removal of others thereby deriving patterns within the structured data, and finally evaluation and interpretation of the output. In this paper a hybrid approach to robustly detect and localize texts in natural scene images is proposed. The variations of text like font, size and line orientation. A text region detector is designed to estimate the text existing confidence and scale information in image pyramid, which helps by segmenting the text components using local binarization. To efficiently filter out the non-text components, a conditional random field (CRF) model considering unary component properties and binary contextual component relationships with supervised parameter learning is proposed. Finally the extracted text components are grouped into text lines/words with a learning-based energy minimization tool and stored in an external file. |
Other Details |
|
Paper ID: IJSRDV2I6293 Published in: Volume : 2, Issue : 6 Publication Date: 01/09/2014 Page(s): 458-461 |
Article Preview |
|
|
|
|
