Alaine Tess A. Cabije
Predicting gender from behavioral features based on large language models (LLMs) is a challenging task with significant implications for personalization, fairness, and privacy. This study evaluates the performance of gender prediction models by examining LLM output patterns, including rethinking assessment and exams (RAE), generic skills (GS) and balanced adoption of artificial intelligence (BA). None of the models showed any meaningful improvement compared to random guessing, considering the validation and test set accuracies were still comparable to chance levels. In particular, GS and BA report the highest correlation (r=0.819, r=0.819), thus suggesting a significant linear relationship between these two variables.However, the variable importance results identify RAE and BA as the most influential predictors in reducing the prediction error, highlighting the need for caution in future research and model development. © 2025 IEEE.
School of Architecture, University of San Carlos, Philippines