Ken D. Gorro, Moustafa F. Ali, Leodivino A. Lawas, Anthony S. Ilano
Natural language processing is a field of computer science that focuses on understanding and analyzing textual data in any given language. Analyzing textual data is very tedious and leads to erroneous results due to unnecessary and noisy data in the corpus. Stop words are considered noisy data which the English language has already predefined corpus of stop words. Stop words in other languages such as Cebuano and Filipino are not yet supported in many NLP API. In the Philippines, users use different languages to post on Facebook. In this study, a corpus of Facebook posts was utilized in automatically detecting a stop word. A neural network was created based on Bidirectional Long Short term memory (BiLSTM). Word2vec was used to provide word embedding and representation from the corpus. The experimental result shows 72% accuracy in using the model. © 2021 ACM.
Department of Industrial Technology, Cebu Technological University, Philippines; Department of Computer, Information Sciences and Mathematics, University of San Carlos, Philippines; Department of Information Technology, Cebu Technological University, Philippines; Department of Fisheries, Cebu Technological University, Philippines