FEATURE SELECTION AND CLASSIFICATION INTEGRATED METHOD FOR IDENTIFYING CITED TEXT SPANS FOR CITANCES ON IMBALANCED DATA

Jen-Yuan Yeh; Cheng-Jung Tsai; Tien-Yu Hsu; Jung-Yi Lin; Pei-Cheng Cheng

doi:10.22452/mjcs.vol34no4.3

FULL TEXT

Published: Oct 31, 2021

DOI: https://doi.org/10.22452/mjcs.vol34no4.3

Keywords:

citation analysis cited text spans identification feature selection classification class imbalance performance evaluation scientific paper summarization

Jen-Yuan Yeh

Dept. of Operation, Visitor Service, Collection and Information Management, National Museum of Natural Science, Taichung 40453, Taiwan

Cheng-Jung Tsai

Graduate Institute of Statistics and Information Science, National Changhua University of Education, Changhua 50007, Taiwan

Tien-Yu Hsu

Dept. of Science Education, National Museum of Natural Science, Taichung 40453, Taiwan

Jung-Yi Lin

IP Affairs Division, Hon Hai Precision Ind.Co.,Ltd (Foxconn), Taipei 11492, Taiwan

Pei-Cheng Cheng

Dept. of Information Management, Chien Hsin University of Science and Technology, Taoyuan 32097, Taiwan

Abstract

Recent studies in scientific paper summarization have explored a new form of structured summary for a reference paper by grouping all cited and citing sentences together by facet. This involves three main tasks: (1) identifying cited text spans for citances (i.e., citing sentences), (2) classifying their discourse facets, and (3) generating a structured summary from the cited text spans and citances. This paper focuses on the first task, and approaches the task as binary classification to distinguish relevant pairs of citances and reference sentences from irrelevant pairs. We propose a new method that integrates feature selection and classification techniques to enhance classification performance. The proposed method investigates combinations of six feature selection methods (χ²-Statistics, Information Gain, Gain Ratio, Relief-F, Significance Attribute Evaluation, and Symmetrical Uncertainty), and five classification algorithms (k-Nearest Neighbors, Decision Tree, Support Vector Machine, Naïve Bayes, and Random Forest). Additionally, to address imbalanced data during training, we apply SMOTE (Synthetic Minority Over-sampling Technique) to introduce synthetic biases towards the minority. Experiments are conducted using the CL-SciSumm corpora to compare the effect of feature selection applied to classification. The results reveal the benefits of feature selection in significantly boosting performance of F₁ score metric, and show that our method is competitive to the state-of-the-art methods in the CL-SciSumm evaluations.

Downloads

How to Cite

Yeh, J.-Y., Tsai, C.-J., Hsu, T.-Y., Lin, J.-Y., & Cheng, P.-C. (2021). FEATURE SELECTION AND CLASSIFICATION INTEGRATED METHOD FOR IDENTIFYING CITED TEXT SPANS FOR CITANCES ON IMBALANCED DATA. Malaysian Journal of Computer Science, 34(4), 355–373. https://doi.org/10.22452/mjcs.vol34no4.3

Issue

Vol. 34 No. 4 (2021): Malaysian Journal of Computer Science

Section

Articles

Article Sidebar

Main Article Content

Abstract

Downloads

Article Details

Most read articles by the same author(s)