Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus
Chengzhi Zhang, Xinyi Yan, Wenqi Yu
Nanjing University of Science and Technology
cs.CL, cs.DL, cs.HC, cs.IR
Submitted: 2026-08-11
Updated: 2026-08-12
Journal ref: aslib JIM, 2026
Code: https://github.com/yan-xinyi/ET_AKE
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
Terminology
Summary
Summary
This paper investigates whether lightweight webcam-based eye-tracking features can enhance keyphrase extraction (KPE) from Chinese academic abstracts in Library and Information Science (LIS). The authors argue that keyphrases are not only statistically or semantically important textual units, but also words or phrases that are more likely to attract readers' attention during comprehension. However, existing KPE studies have mainly focused on improving models' representation learning from input texts while largely overlooking the connection between keyphrases and human reading behavior.
Methodology/Approach: Motivated by the limited availability of eye-tracking data for Chinese academic reading, the authors developed a lightweight webcam-based data collection platform by integrating the open-source SearchGazer library. They constructed the Chinese LIS Eye-Tracking Corpus (CLIS-ET) based on collected and preprocessed data. The study incorporated three character-level eye-tracking features—first fixation duration (FFD), fixation number (FN), and total fixation duration (TFD)—into KPE models to evaluate their effects on extraction performance. The data collection experiment involved ten graduate students or senior undergraduates from the LIS field. The reading materials consisted of 320 article titles and abstracts (Abstract-320) for the eye-tracking experiment, and a larger dataset of 5,190 articles (Abstract-5190) for the keyphrase extraction task. The KPE models tested included recurrent neural network-based models (BiLSTM, BiLSTM+CRF, Att-BiLSTM, Att-BiLSTM+CRF) and pre-trained language model-based models (BERT, RoBERTa, MacBERT, and Randeng-T5).
Findings: By integrating eye-tracking features into KPE models, the experiments demonstrated that the combination of fixation number and total fixation duration (FN+TFD) achieved the best performance on the Att-BiLSTM+CRF-based KPE model, highlighting the significant impact of eye-tracking data on optimizing keyphrase extraction from academic literatures. The results showed that eye-tracking features provide complementary cognitive signals for KPE. On the smaller Abstract-320 dataset, FFD achieved the most pronounced improvement on the BERT model, reaching an enhancement of 2.83%. On the larger Abstract-5190 dataset, RoBERTa attained its optimal performance of 44.51% when combined with FN. The significance analysis using two-tailed paired t-tests confirmed that eye-tracking features exert statistically significant effects across multiple KPE architectures, with FN+TFD demonstrating the most consistent significance across model types.
Originality/Value: This paper presents a cost-effective eye-tracking methodology to improve keyphrase extraction from academic papers. The authors introduce the Chinese Academic Eye-Tracking Corpus (CLIS-ET), which includes key eye-tracking metrics: fixation number, first fixation duration, and total fixation duration. The findings demonstrate that incorporating these features enhances KPE models. Compared with traditional eye-tracking methods that rely on expensive equipment and controlled laboratory environments, the proposed approach uses ordinary webcams and open-source frameworks, substantially reducing experimental and deployment costs while maintaining reliable fixation accuracy. The developed data collection framework supports large-scale, synchronized eye-tracking acquisition and can serve as a foundation for cognitive NLP research in future extensions to other academic domains and participant groups. The dataset and source code can be accessed at: https://github.com/yan-xinyi/ET AKE.
Improvements for AI systems
Improvements to AI Systems:
- Cognitive-Aware Keyphrase Extraction (KPE) Models
-
Integrate real-time eye-tracking features (FFD, FN, TFD) as auxiliary inputs into existing KPE architectures (e.g., BiLSTM+CRF, BERT, RoBERTa).
-
The improved system can predict keyphrases that align with human reading attention, not just statistical salience, leading to higher relevance for summarization, indexing, and retrieval tasks.
- Lightweight, Webcam-Based Cognitive Signal Collection
-
Embed the SearchGazer-based data collection framework into AI pipelines to capture gaze data without specialized hardware.
-
The improved system can self-calibrate and adapt to individual users’ reading patterns, enabling personalized keyphrase extraction in real-time for applications like e-readers, research assistants, or document summarization tools.
- Cross-Domain Transferable Cognitive Features
-
Use the CLIS-ET corpus to pre-train a feature extractor that maps gaze signals to textual importance weights.
-
The improved system can transfer these learned cognitive priors to other languages or domains (e.g., English, medical, legal texts) where eye-tracking data is scarce, boosting KPE performance without additional data collection.
- Adaptive Feature Selection for Model Optimization
-
Implement a dynamic feature-weighting mechanism that selects the best eye-tracking feature combination (e.g., FN+TFD vs. FFD) based on dataset size and model architecture.
-
The improved system can automatically tune its cognitive input strategy, achieving up to 2.83% (BERT, small data) and 44.51% (RoBERTa, large data) performance gains, while maintaining robustness across different model families.
- Cognitive Signal Fusion for Pre-trained Language Models
-
Add a gaze-attention layer that fuses eye-tracking features with transformer-based embeddings, allowing the model to attend to words that humans fixate on longer.
-
The improved system can generate more human-like keyphrases for academic abstracts, improving downstream tasks like automatic literature review, research trend analysis, and citation recommendation.
- Low-Cost, Scalable Cognitive NLP Benchmarking
-
Use the CLIS-ET corpus and the open-source framework to evaluate new KPE models against cognitive baselines.
-
The improved system can serve as a standardized testbed for cognitive-aware NLP, enabling fair comparison of models that incorporate human reading behavior, and accelerating research in cognitive computing and human-AI interaction.
Abstract
Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking human reading behavior. This study examines whether lightweight webcam-based eye-tracking features can improve KPE from Chinese academic abstracts in Library and Information Science (LIS). Methodology: To address the limited availability of eye-tracking data for Chinese academic reading, we developed a lightweight webcam-based data collection platform using the open-source SearchGazer library and constructed the Chinese LIS Eye-Tracking Corpus (CLIS-ET). Three character-level eye-tracking features, first fixation duration (FFD), fixation number (FN), and total fixation duration (TFD), were incorporated into KPE models to evaluate their effects on extraction performance. Findings: Eye-tracking features consistently improved KPE performance. The combination of FN and TFD achieved the best results on the Att-BiLSTM+CRF model, indicating that readers' fixation behavior provides useful signals for identifying keyphrases in academic abstracts. Originality/value: This study introduces a cost-effective webcam-based eye-tracking approach for KPE and presents CLIS-ET, a Chinese academic eye-tracking corpus containing FFD, FN, and TFD features. The results demonstrate the value of incorporating human reading behavior into keyphrase extraction. Dataset and code: https://github.com/yan-xinyi/ET AKE and https://github.com/yan-xinyi/Reading ET System.
Sources
- Advancing NLP with Cognitive Language Processing Signals
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering