Optimization of k-Nearest Neighbour to categorize Indonesian’s news articles

Ihsan, Afdhalul and Rainarli, Ednawati (2021) Optimization of k-Nearest Neighbour to categorize Indonesian’s news articles. Asia-Pacific Journal of Information Technology and Multimedia, 10 (1). pp. 43-51. ISSN 2289-2192

[img]
Preview
PDF
259kB

Official URL: https://www.ukm.my/apjitm/articles-year.php

Abstract

Text classification is the process of grouping documents based on similarity in categories. Some of the obstacles in doing text classification are many words appeared in the text, and some words come up with infrequent frequency (sparse words). The way to solve this problem is to conduct the feature selection process. There are several filter-based feature selection methods; some are Chi-Square, Information Gain, Genetic Algorithm, and Particle Swarm Optimization (PSO). Aghdam's research shows that PSO is the best among those methods. This study examined PSO to optimize the k-Nearest Neighbour (k-NN) algorithm's performance in categorizing news articles. k-NN is an algorithm that is simple and easy to implement. If we use the appropriate features, then the k-NN will be a reliable algorithm. PSO algorithm is used to select keywords (term features), and it is continued with classifying the documents using k-NN. The testing process consists of three stages. The stages are tuning the parameter of k-NN, the parameter of PSO, and measuring the testing performance. The parameter tuning process aims to determine the number of neighbours used in k-NN and optimize the PSO particles. Otherwise, the performance testing compares the performance of k-NN with and without using PSO. The optimal number of neighbours is 9, with the number of particles is 50. The testing showed that using the k-NN with PSO and a 50% reduction in terms. The results 20 per cent better accuracy than k-NN without PSO. Although the PSO's process did not always find the optimal conditions, the k-NN method can produce better accuracy. In this way, the k-NN method can work better in grouping news articles, especially in Indonesian language news articles.

Item Type:Article
Keywords:Feature selection; k-Nearest Neighbour; Metaheuristic; Optimization; Text classification
Journal:Asia - Pasific Journal of Information Technology and Multimedia (Formerly Jurnal Teknologi Maklumat dan Multimedia)
ID Code:16843
Deposited By: ms aida -
Deposited On:15 Jun 2021 03:42
Last Modified:20 Jun 2021 05:01

Repository Staff Only: item control page