WAYS TO DETERMINE THE RANGE OF KEYWORDS IN A FREQUENCY DICTIONARY FOR TEXT CLASSIFICATION

Authors

  • Olesia Barkovska Kharkiv National University of Radio Electronics https://orcid.org/0000-0001-7496-4353
  • Dmytro Mohylevskyi Kharkiv National University of Radio Electronics
  • Yuliia Ivanenko Kharkiv national university of radio electronics
  • Dmytro Rosinskiy Kharkiv national university of radio electronics

DOI:

https://doi.org/10.31891/csit-2023-1-2

Keywords:

processing, language, vocabulary, frequency, term, keyword, сlassification, feature

Abstract

The paper is devoted to the actual problem of classifying textual documents of the collection by characteristic features, which is used for classifying news, reviews, determining the emotional tone of the text, as well as for forming catalogs of scientific, academic and research works. The paper proposes an approach for determining the significant words of a document for their further use as a feature vector in the classification process. In the course of the work, the author's keywords were identified, a partial dictionary was built, and the correlation between the author's keywords and the list of ordered words of the frequency dictionary based on the TF method, which also includes the author’s keywords, was analyzed. The determination of the range and percentage of significant words allows for further classification of scientific and research papers when forming thematic catalogs even in the absence of a list of author's keywords that can be used for classification. The results show that the use of the entire input range of frequency dictionary words is redundant and leads to a longer classification time.

Downloads

Published

2023-03-30

How to Cite

Barkovska, O., Mohylevskyi , D., Ivanenko , Y., & Rosinskiy, D. (2023). WAYS TO DETERMINE THE RANGE OF KEYWORDS IN A FREQUENCY DICTIONARY FOR TEXT CLASSIFICATION. Computer Systems and Information Technologies, (1), 14–20. https://doi.org/10.31891/csit-2023-1-2