Text normalization method for arabic handwritten script

Tarik Abu-Ain, Siti Norul Huda Sheikh Abdullah, Khairuddin Omar, Ashraf Abu-Ein, Bilal Bataineh, Waleed Abu-Ain

Research output: Contribution to journalArticle

2 Citations (Scopus)

Abstract

Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line. The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance.

Original languageEnglish
Pages (from-to)164-175
Number of pages12
JournalJournal of ICT Research and Applications
Volume7
Issue number2
DOIs
Publication statusPublished - 2013

Fingerprint

Straightening
Image recognition
Character recognition
Image analysis
Feature extraction
Experiments
Normalization

Keywords

  • Arabic handwriting
  • Baseline detection
  • Handwritten text normalization
  • Preprocessing
  • Slant correction
  • Sub-word extraction

ASJC Scopus subject areas

  • Computer Science(all)
  • Electrical and Electronic Engineering
  • Information Systems and Management

Cite this

Text normalization method for arabic handwritten script. / Abu-Ain, Tarik; Sheikh Abdullah, Siti Norul Huda; Omar, Khairuddin; Abu-Ein, Ashraf; Bataineh, Bilal; Abu-Ain, Waleed.

In: Journal of ICT Research and Applications, Vol. 7, No. 2, 2013, p. 164-175.

Research output: Contribution to journalArticle

Abu-Ain, Tarik ; Sheikh Abdullah, Siti Norul Huda ; Omar, Khairuddin ; Abu-Ein, Ashraf ; Bataineh, Bilal ; Abu-Ain, Waleed. / Text normalization method for arabic handwritten script. In: Journal of ICT Research and Applications. 2013 ; Vol. 7, No. 2. pp. 164-175.
@article{ef852488dd60418c89247c13060cc225,
title = "Text normalization method for arabic handwritten script",
abstract = "Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line. The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance.",
keywords = "Arabic handwriting, Baseline detection, Handwritten text normalization, Preprocessing, Slant correction, Sub-word extraction",
author = "Tarik Abu-Ain and {Sheikh Abdullah}, {Siti Norul Huda} and Khairuddin Omar and Ashraf Abu-Ein and Bilal Bataineh and Waleed Abu-Ain",
year = "2013",
doi = "10.5614/itbj.ict.res.appl.2013.7.2.5",
language = "English",
volume = "7",
pages = "164--175",
journal = "Journal of ICT Research and Applications",
issn = "2337-5787",
publisher = "Institut Teknologi Bandung (ITB)",
number = "2",

}

TY - JOUR

T1 - Text normalization method for arabic handwritten script

AU - Abu-Ain, Tarik

AU - Sheikh Abdullah, Siti Norul Huda

AU - Omar, Khairuddin

AU - Abu-Ein, Ashraf

AU - Bataineh, Bilal

AU - Abu-Ain, Waleed

PY - 2013

Y1 - 2013

N2 - Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line. The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance.

AB - Text normalization is an important technique in document image analysis and recognition. It consists of many preprocessing stages, which include slope correction, text padding, skew correction, and straight the writing line. In this side, text normalization has an important role in many procedures such as text segmentation, feature extraction and characters recognition. In the present article, a new method for text baseline detection, straightening, and slant correction for Arabic handwritten texts is proposed. The method comprises a set of sequential steps: first components segmentation is done followed by components text thinning; then, the direction features of the skeletons are extracted, and the candidate baseline regions are determined. After that, selection of the correct baseline region is done, and finally, the baselines of all components are aligned with the writing line. The experiments are conducted on IFN/ENIT benchmark Arabic dataset. The results show that the proposed method has a promising and encouraging performance.

KW - Arabic handwriting

KW - Baseline detection

KW - Handwritten text normalization

KW - Preprocessing

KW - Slant correction

KW - Sub-word extraction

UR - http://www.scopus.com/inward/record.url?scp=84901789644&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84901789644&partnerID=8YFLogxK

U2 - 10.5614/itbj.ict.res.appl.2013.7.2.5

DO - 10.5614/itbj.ict.res.appl.2013.7.2.5

M3 - Article

VL - 7

SP - 164

EP - 175

JO - Journal of ICT Research and Applications

JF - Journal of ICT Research and Applications

SN - 2337-5787

IS - 2

ER -