Text clustering for reducing semantic information in Malay semantic representation

Tuan Norhafizah Tuan Zakaria, and Mohd Juzaiddin Ab Aziz, and Mohd Rosmadi Mokhtar, and Saadiyah Darus, (2020) Text clustering for reducing semantic information in Malay semantic representation. Asia-Pacific Journal of Information Technology and Multimedia, 9 (2). pp. 11-24. ISSN 2289-2192

[img]
Preview
PDF
457kB

Official URL: https://www.ukm.my/apjitm/articles-year.php

Abstract

The generation of texts are dramatically increased in this era. A text basically consists of structured and unstructured texts. The enormous amount of unstructured texts can be easily perceived by humans, unfortunately cannot be simply processed by computer. It needs efficient techniques to reduce the information into more valuable vectors. In this article, we introduce text clustering method using Malay linguistic information to reduce the unstructured semantic information derived from Wikipedia Bahasa Melayu’s articles. The proposed method uses the linguistic features in Malay language to cater the morphological issues of Malay words. We have incorporated semantic information from semantic lexical resource for Malay, which called Wikipedia Bahasa Melayu (WikiBM). Then, an experiment was conducted to evaluate the effects of text clustering to the semantic similarity value using gloss definition of WikiBM’s article. We used Jaccard similarity to calculate the overlaps vectors from the text of WikiBM. Then, the correlation was computed using Pearson’s correlation. The score between original text definition was compared to the new text definition using text clustering method. From the experiment, we can conclude that the correlation value was increased after the semantic information was reduced to more valuable vectors using text clustering method (from 0.39 to 0.43).

Item Type:Article
Keywords:Text clustering; Malay; Semantic representation; Wikipedia Bahasa Melayu; Semantic similarity measurement
Journal:Asia - Pasific Journal of Information Technology and Multimedia (Formerly Jurnal Teknologi Maklumat dan Multimedia)
ID Code:16833
Deposited By: ms aida -
Deposited On:15 Jun 2021 00:41
Last Modified:20 Jun 2021 04:35

Repository Staff Only: item control page