Abstract
Dyslexia is a specific learning difficulty associated with persistent difficulties in accurate and fluent word recognition, decoding and spelling. Early identification can facilitate timely educational intervention, yet conventional assessment may be resource-intensive and machine-learning prediction is complicated by high-dimensional behavioural data and substantial class imbalance. This study develops and empirically evaluates an ensemble machine-learning framework for early dyslexia prediction using Random-Forest Recursive Feature Elimination (RF-RFE) and heterogeneous tree-based learners. The empirical analysis uses the Dyt-desktop behavioural dataset containing 3,644 observations, 196 predictor variables and 392 dyslexia-positive cases. The dataset represents a strongly imbalanced binary classification problem, with dyslexia cases accounting for 10.76% of observations. The proposed framework integrates data preprocessing, leakage-controlled RF-RFE feature selection, Random Forest (RF), Extreme Gradient Boosting (XGBoost), Extra Trees (ET), and probability-level stacking. Stratified five-fold out-of-fold validation was used to evaluate predictive performance. Accuracy, precision, recall, specificity, F1-score, receiver operating characteristic area under the curve (ROC-AUC), and Matthews correlation coefficient (MCC) were used as evaluation measures. The baseline XGBoost model achieved the strongest individual performance, obtaining 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE combined with XGBoost improved performance to 90.81% accuracy, 64.77% precision, 31.89% recall, 97.91% specificity, 42.74% F1-score, 0.8858 ROC-AUC and 0.4122 MCC. The heterogeneous RF-XGBoost-ET stacking ensemble achieved the highest recall of 71.17%, F1-score of 51.71% and MCC of 0.4644, although accuracy declined to 85.70%. The findings demonstrate that recursive feature elimination can improve the predictive representation of high-dimensional behavioural data, particularly when combined with boosting. The results also demonstrate that ensemble learning provides a substantially more sensitivity-oriented operating point than individual classifiers. Because of class imbalance, accuracy alone is shown to be inadequate for evaluating dyslexia screening models. The study contributes a reproducible RF-RFE ensemble framework for AI-assisted early dyslexia screening and recommends threshold optimization, precision-recall analysis, explainability, fairness assessment and external validation before deployment.