Joshua Patrick G. Chiu, Kyne Edric B. Sia, Christine D. Bandalan
The Philippines, having over 170 languages, faces language barriers and diminished language fluency due to heavy English influence and the effects of globalization. The language focuses on Cebuano, the second most spoken language mainly spoken in Central Visayas. Existing Cebuano parsers have been made through Recursive Descent Parsing (RDP) for a Cebuano Parse Tree in relaying grammar through a syntactic structure. However, there were identified limitations through low-level sentences, lacking grammar codes, low metric score, and stemmer evaluation errors. To address this gap, the researchers proposed a Snowball-based stemmer approach along with the Cocke-Kasami-Younger algorithm and additional grammar codes for better syntax tree optimization. Due to the lack of a dataset for Cebuano sentences, the Cebuano Corpus is extensively created by the researchers via web scraping and data augmentation techniques such as language dialect filtration and sentence range variation. The dataset comprises at least 500 collected phrases to be annotated by the linguist experts for the Gold Standard Parse Treebank (GSPT). The GSPT will serve as the parser's benchmark through syntax constituents, known as the PARSEVAL metric system. The improved parser attained an accuracy of 82.14%, suggesting a high rate of correct Cebuano constituents from a Cebuano sentence. The findings suggest that the developed system is more effective than its predecessor. By leveraging advanced Snowball-based algorithms and updated Grammar Codes, the study demonstrates the potential of parsers in preserving cultural language. © 2025 IEEE.
University of San Carlos, Information Sciences and Mathematics, Department of Computer, Cebu, Philippines