Person-independent and cross-dataset evaluation of convolutional and recurrent architectures for eye movement classification
Classifying eye movements into fixations, saccades, and smooth pursuits is a basic step in the analysis of gaze recordings. Deep models are now common for this task, but they are often evaluated on windows that share participants between training and test sets, which leaves open how well they work for people the model has not seen. This paper reports a person-independent comparison of three architectures—a one-dimensional CNN, an LSTM, and a CNN–LSTM hybrid—for four-class eye movement classification on the GazeCom dataset (54 participants, 250 Hz, expert frame-level annotations).
Each model is trained and tested under leave-one-person-out cross-validation, so every participant is held out in turn, and each architecture is run with two input representations: raw gaze position and a derived set that adds velocity, acceleration, and angular change. Model differences are assessed at the participant level (paired tests over 54 folds, Holm-corrected) rather than on overlapping windows. Two findings are consistent across the study.
First, recurrent modeling is what separates the strong models from the weak one: the LSTM and the hybrid reach a macro-F1 of about 0.82 with the derived input, while the CNN reaches 0.71, and the gap holds for every participant. Second, adding the convolutional front end to the LSTM does not help: the hybrid and the plain LSTM are statistically indistinguishable with the raw input and separated by less than 0.01 macro-F1 with the derived input, so the extra component is not warranted by the data; the hybrid carries about 70% more parameters than the LSTM for no measurable gain. The derived input improves every architecture, most for the CNN and for the smooth-pursuit class.
Three non-learned or lightweight baselines (velocity threshold, velocity–dispersion, and a Random Forest on kinematic features), tuned per fold and evaluated under the same protocol, stay near 0.53–0.54 macro-F1 on the fixation/saccade/pursuit task, well below all three deep models. When the GazeCom-trained models are tested without retraining on the Lund2013 recordings and scored against both expert raters, accuracy falls as expected, but the ordering of the architectures is preserved. The plain LSTM is therefore an adequate model for this task; the smooth-pursuit class remains the main source of error and is the most useful target for future work.
During the preparation of this manuscript, the author used Paperpal to support language editing and to improve the fluency and clarity of the English text. The author reviewed and edited all generated content and takes full responsibility for the final manuscript. Department of Computer Technologies, Bandırma Onyedi Eylül University, Bandırma, 10900, Gönen, Balıkesir, Turkey The authors declare no competing interests.
This study used only publicly available, de-identified eye movement datasets (GazeCom and Lund2013) and did not involve any new data collection from human participants. Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Extract — continue reading at the source.