Text complexity and linguistic features: Their correlation in English and Russian
Journal: Russian Journal of Linguistics (Vol.26, No. 2)Publication Date: 2022-06-30
Authors : Dmitry Morozov; Anna Glazkova; Boris Iomdin;
Page : 426-448
Keywords : text complexity; machine learning; neural network; corpus linguistics;
Abstract
Text complexity assessment is a challenging task requiring various linguistic aspects to be taken into consideration. The complexity level of the text should correspond to the reader’s competence. A too complicated text could be incomprehensible, whereas a too simple one could be boring. For many years, simple features were used to assess readability, e.g. average length of words and sentences or vocabulary variety. Thanks to the development of natural language processing methods, the set of text parameters used for evaluating readability has expanded significantly. In recent years, many articles have been published the authors of which investigated the contribution of various lexical, morphological, and syntactic features to the readability level. Nevertheless, as the methods and corpora are quite diverse, it may be hard to draw general conclusions as to the effectiveness of linguistic information for evaluating text complexity due to the diversity of methods and corpora. Moreover, a cross-lingual impact of different features on various datasets has not been investigated. The purpose of this study is to conduct a large-scale comparison of features of different nature. We experimentally assessed seven commonly used feature types (readability, traditional features, morphological features, punctuation, syntax frequency, and topic modeling) on six corpora for text complexity assessment in English and Russian employing four common machine learning models: logistic regression, random forest, convolutional neural network and feedforward neural network. One of the corpora, the corpus of fiction literature read by Russian school students, was constructed for the experiment using a large-scale survey to ensure the objectivity of the labeling. We showed which feature types can significantly improve the performance and analyzed their impact according to the dataset characteristics, language, and data source.
Other Latest Articles
- Collection and evaluation of lexical complexity data for Russian language using crowdsourcing
- A cognitive linguistic approach to analysis and correction of orthographic errors
- What neural networks know about linguistic complexity
- ReaderBench: Multilevel analysis of Russian text characteristics
- Natural language processing and discourse complexity studies
Last modified: 2022-06-30 03:46:24