상세 보기
Evaluation of Text Classification Methods for Small Labeled Korean Datasets Using Multiple-Criteria Decision-making Methods
- Lee Minkyu;
- Kim Namhyoung
WEB OF SCIENCE
1SCOPUS
1초록
Preprocessing and feature selection are important steps in the machine learning pipeline for text classification. These steps ensure that the input data are appropriately prepared for model training and that the most informative features are selected for the model to learn from. Morpheme analyzer is also crucial in machine learning performance, particularly in text classification tasks that involve languages with complex morphology. This study compares the performance of several feature selection methods, Korean morpheme analyzers, and classifiers in small-scale Korean text classification. PROMETHEE, a multi-criteria decision-making method (MCDM), is used to rank each method based on its performance. As per the results, the document frequency method exhibited the best overall performance among several feature selection methods. In terms of Korean morpheme analyzers, the most recently developed Khaiii demonstrated the best performance. The support vector machine classifier also exhibits suitable performance, suggesting that it is a reliable choice for small-scale Korean text classification. While performance differences among the methods were not significant for each dataset, the overall trend was clear, the aforementioned 3 methods consistently outperformed the others. Through this study, identification of effective techniques for small-scale Korean text classification is feasible, and it is expected to provide guidelines to analysists.
키워드
- 제목
- Evaluation of Text Classification Methods for Small Labeled Korean Datasets Using Multiple-Criteria Decision-making Methods
- 저자
- Lee Minkyu; Kim Namhyoung
- 발행일
- 2024-09
- 유형
- Article
- 권
- 23
- 호
- 3
- 페이지
- 311 ~ 322