Roxanne Angelli P. Lopez, Ernie R. Manatad, Christine D. Bandalan
In recent years, Sign Language Recognition (SLR) has grown with advancements in computer vision, machine learning, and deep learning. Studies have explored a variety of feature extraction models, alongside sequence classifiers to effectively process spatial and temporal information in both static and dynamic gestures. Although these models have been able to achieve high accuracy independently, each model has its own strengths and weaknesses. In this study, the effect of a soft-voting ensemble technique was investigated on classification performance within a controlled experimental setting. Multiple model configurations were trained and evaluated under identical conditions, with the highest performing selected for integration into the ensemble system. The approach was ultimately able to improve overall performance, achieving an increase ranging from 0.75% to 1.38% in accuracy, precision, recall, and F1 score when compared to the performance results of the highest performing models individually. © 2026 IEEE.
University of San Carlos (USC), Department of Computer, Information Science, and Mathematics (DCISM), Cebu City, Philippines