Bootstrapping Dependency Treebank of Urdu Noisy Text
Journal: International Journal of Emerging Trends in Engineering Research (IJETER) (Vol.9, No. 8)Publication Date: 2021-08-12
Authors : Amber Baig Mutee U Rahman Sirajuddin Qureshi Saima Tunio Sehrish Abrejo Shadia S. Baloch;
Page : 1102-1106
Keywords : ;
Abstract
Amber Baig et al., International Journal of Emerging Trends in Engineering Research, 9(8), August 2021, 1102 – 1106 1102 ABSTRACT This paper describes how bootstrapping was used to extend the development of the Urdu Noisy Text dependency treebank. To overcome the bottleneck of manually annotating corpus for a new domain of user-generated text, MaltParser, an opensource, data-driven dependency parser, is used to bootstrap the treebank in semi-automatic manner for corpus annotation after being trained on 500 tweet Urdu Noisy Text Dependency Treebank. Total four bootstrapping iterations were performed. At the end of each iteration, 300 Urdu tweets were automatically tagged, and the performance of parser model was evaluated against the development set. 75 automatically tagged tweets were randomly selected out of pre-tagged 300 tweets for manual correction, which were then added in the training set for parser retraining. Finally, at the end of last iteration, parser performance was evaluated against test set. The final supervised bootstrapping model obtains a LA of 72.1%, UAS of 75.7% and LAS of 64.9%, which is a significant improvement over baseline score of 69.8% LA, 74% UAS, and 62.9% LAS
Other Latest Articles
- A Review of Skin Disease Classification Techniques based on Machine Learning
- A Survey on Trending Topics of Microservices
- Modification and Performance Test of Prototype I Washer Machine
- Modelling and Performance Analysis of a DSTATCOM using ISCT Control Technique
- A Novel Kernelized Fuzzy Clustering Algorithm for Data Classification
Last modified: 2021-08-13 00:38:06