Automating Rey Complex Figure Test scoring using
a deep learning-based approach: A potential large-
scale screening tool for congnitive decline
Jun Young Park
Seoul National University
Eun Hyun Seo
Gwangju Alzheimer’s & Related Dementia cohort research center, Chosun University
Hyung-Jun Yoon
Chosun University
Sungho Won
Seoul National University
Kun Ho Lee ( )
Gwangju Alzheimer’s & Related Dementia cohort research center, Chosun University
Research Article
Keywords: Alzheimer’s Disease, Rey Complex Figure Test, Scoring, Arti cial Intelligence, Deep Learning,
Convolutional Neural Network.
Posted Date: August 23rd, 2022
DOI: https://doi.org/10.21203/rs.3.rs-1973305/v1
License: This work is licensed under a Creative Commons Attribution 4.0 International License.
Read Full License
Page 1/20
,Abstract
Background:
The Rey Complex Figure Test (RCFT) has been widely used to evaluate neurocognitive functions in
various clinical groups with a broad range of ages. However, despite its usefulness, the scoring method is
as complex as the gure. Such a complicated scoring system can lead to the risk of reducing the extent
of agreement among raters. Although several attempts have been made to use RCFT in clinical settings
in a digitalized format, little attention has been given to develop direct automatic scoring that is
comparable to experienced psychologists. Therefore, we aimed to develop an arti cial intelligence (AI)
scoring system for RCFT using a deep learning (DL) algorithm and con rmed its validity.
Methods:
A total of 6,680 subjects were enrolled in the Gwangju Alzheimer’s and Related Dementia cohort registry,
Korea from January 2015 to June 2021. We obtained 20,040 scanned images using three images per
subject (copy, immediate recall, and delayed recall) and scores rated by 32 experienced psychologists. We
trained the automated scoring system using the DenseNet architecture. To increase the model
performance, we improved the quality of training data by re-examining some images with poor results
(mean absolute error (MAE) 5 [points]) and re-trained our model. Finally, we conducted an external
validation with 150 images scored by ve experienced psychologists.
Results:
For ve-fold cross-validation, our rst model obtained MAE = 1.24 [points] and R-squared ( ) = 0.977.
However, after evaluating and updating the model, the performance of the nal model was improved
(MAE = 0.95 [points], = 0.986). Predicted scores among cognitively normal, mild cognitive impairment,
and dementia were signi cantly differed. For the 150 independent test sets, the MAE and between AI and
average scores by ve human experts was 0.64 [points] and 0.994, respectively.
Conclusion:
We concluded that there was no fundamental difference between the rating scores of experienced
psychologists and those of our AI scoring system. We expect that our AI psychologist will be able to
contribute to screen the early stages of Alzheimer’s disease pathology in medical checkup centers or
large-scale community-based research institutes in a faster and cost-effective way.
Background
The Rey Complex Figure Test (RCFT) was originally developed to evaluate the perceptual organization
and visual memory [1]. It is valuable and practical in that the test is relatively simple and clear to
administer, and it assesses multiple cognitive domains, including executive function and visuospatial
ability or memory. [1, 2]. The RCFT has been widely used to evaluate neurocognitive functions in various
Page 2/20
, clinical groups with a broad range of ages [3, 4]. Visuospatial modality of episodic memory has been
suggested as having a signi cant association with tau pathology in Alzheimer’s disease (AD) [5–8].
Particularly, previous studies on Alzheimer’s continuum have demonstrated that RCFT scores can be an
early marker of clinical progression [9] or tau pathology [5]. In addition, the RCFT sensitively captures
organizational strategies in healthy young adults [10] and in patients with brain damage [11, 12].
Several quantitative and qualitative scoring systems have been proposed [3]. The most broadly used
method is the 18-item and 36-point scoring system standardized by Osterrieth [13]. However, despite its
usefulness, the scoring method is as complex as the gure. Complex scoring system could lead to the
risk of reducing the extent of agreement among raters [14]. Therefore, it is essential to acuqire scoring
skills before administration of the RCFT. Raters for the RCFT need to be trained intensively to score
reliably. Consequently, conducting the RCFT in large-scale community-based studies is di cult. However,
the demand for digital-based cognitive assessments has increased. The digitalization of cognitive
assessment has developed rapidly with technological advancements [15]. Traditional cognitive measures
such as the RCFT are reliable candidates for digitalization. Establishing an automatic scoring system for
the RCFT could be an unavoidable initial step in the evolution of digital cognitive assessments.
Recently, as deep learning (DL) has undergone remarkable improvements in health care [16], there have
been some efforts toward automating the assessment of digitalized drawing tests such as the pentagon
drawing test (PDT) and the clock drawing test (CDT). In particular, convolutional neural networks (CNN)
that can extract important features automatically from raw data have been widely used and shown
excellent performance. Several automatic scoring systems for PDT were developed using CNN [17–19].
Previous studies on digitalized CDT have shown that CNN could distinguish subjects with cognitive
impairment from cognitively normal (CN) subjects [4, 20, 21].
Meanwhile, for digitalized RCFT, many studies using computer vision technology have been proposed.
For instance, digital tablets and pens have been widely used to generate images and extract distinctive
features from digitalized images. [22] showed the differences between adolescents with attention-de cit
hyperactivity disorder and healthy adolescents by comparing the pixel mean between a template image
and images drawn using a digital tablet. Also, the pen stroke data and spatial information from images
drawn by a digital tablet and pen were also used to distinguish subjects with AD from CN subjects [23].
Furthermore, several DL methods were proposed, recently for the digitalized RCFT. CNN methods using
raw RCFT images have been applied to differentiate individuals with cognitive impairment from those
with CN [24–26]. Those studies have focused on identifying the different patterns between clinical
diagnostic groups and healthy controls.
However, methods for directly predicting RCFT scores based on the 36-point scoring system, which is
widely used in clinical elds, have been very limited. A method to score the RCFT was rstly developed by
segmenting six relevant scoring sections [27]. However, it offered only six of the 18 scoring sections, so
could not be applied to the 36-point scoring system. A DL method for scoring the RCFT was proposed
Page 3/20
a deep learning-based approach: A potential large-
scale screening tool for congnitive decline
Jun Young Park
Seoul National University
Eun Hyun Seo
Gwangju Alzheimer’s & Related Dementia cohort research center, Chosun University
Hyung-Jun Yoon
Chosun University
Sungho Won
Seoul National University
Kun Ho Lee ( )
Gwangju Alzheimer’s & Related Dementia cohort research center, Chosun University
Research Article
Keywords: Alzheimer’s Disease, Rey Complex Figure Test, Scoring, Arti cial Intelligence, Deep Learning,
Convolutional Neural Network.
Posted Date: August 23rd, 2022
DOI: https://doi.org/10.21203/rs.3.rs-1973305/v1
License: This work is licensed under a Creative Commons Attribution 4.0 International License.
Read Full License
Page 1/20
,Abstract
Background:
The Rey Complex Figure Test (RCFT) has been widely used to evaluate neurocognitive functions in
various clinical groups with a broad range of ages. However, despite its usefulness, the scoring method is
as complex as the gure. Such a complicated scoring system can lead to the risk of reducing the extent
of agreement among raters. Although several attempts have been made to use RCFT in clinical settings
in a digitalized format, little attention has been given to develop direct automatic scoring that is
comparable to experienced psychologists. Therefore, we aimed to develop an arti cial intelligence (AI)
scoring system for RCFT using a deep learning (DL) algorithm and con rmed its validity.
Methods:
A total of 6,680 subjects were enrolled in the Gwangju Alzheimer’s and Related Dementia cohort registry,
Korea from January 2015 to June 2021. We obtained 20,040 scanned images using three images per
subject (copy, immediate recall, and delayed recall) and scores rated by 32 experienced psychologists. We
trained the automated scoring system using the DenseNet architecture. To increase the model
performance, we improved the quality of training data by re-examining some images with poor results
(mean absolute error (MAE) 5 [points]) and re-trained our model. Finally, we conducted an external
validation with 150 images scored by ve experienced psychologists.
Results:
For ve-fold cross-validation, our rst model obtained MAE = 1.24 [points] and R-squared ( ) = 0.977.
However, after evaluating and updating the model, the performance of the nal model was improved
(MAE = 0.95 [points], = 0.986). Predicted scores among cognitively normal, mild cognitive impairment,
and dementia were signi cantly differed. For the 150 independent test sets, the MAE and between AI and
average scores by ve human experts was 0.64 [points] and 0.994, respectively.
Conclusion:
We concluded that there was no fundamental difference between the rating scores of experienced
psychologists and those of our AI scoring system. We expect that our AI psychologist will be able to
contribute to screen the early stages of Alzheimer’s disease pathology in medical checkup centers or
large-scale community-based research institutes in a faster and cost-effective way.
Background
The Rey Complex Figure Test (RCFT) was originally developed to evaluate the perceptual organization
and visual memory [1]. It is valuable and practical in that the test is relatively simple and clear to
administer, and it assesses multiple cognitive domains, including executive function and visuospatial
ability or memory. [1, 2]. The RCFT has been widely used to evaluate neurocognitive functions in various
Page 2/20
, clinical groups with a broad range of ages [3, 4]. Visuospatial modality of episodic memory has been
suggested as having a signi cant association with tau pathology in Alzheimer’s disease (AD) [5–8].
Particularly, previous studies on Alzheimer’s continuum have demonstrated that RCFT scores can be an
early marker of clinical progression [9] or tau pathology [5]. In addition, the RCFT sensitively captures
organizational strategies in healthy young adults [10] and in patients with brain damage [11, 12].
Several quantitative and qualitative scoring systems have been proposed [3]. The most broadly used
method is the 18-item and 36-point scoring system standardized by Osterrieth [13]. However, despite its
usefulness, the scoring method is as complex as the gure. Complex scoring system could lead to the
risk of reducing the extent of agreement among raters [14]. Therefore, it is essential to acuqire scoring
skills before administration of the RCFT. Raters for the RCFT need to be trained intensively to score
reliably. Consequently, conducting the RCFT in large-scale community-based studies is di cult. However,
the demand for digital-based cognitive assessments has increased. The digitalization of cognitive
assessment has developed rapidly with technological advancements [15]. Traditional cognitive measures
such as the RCFT are reliable candidates for digitalization. Establishing an automatic scoring system for
the RCFT could be an unavoidable initial step in the evolution of digital cognitive assessments.
Recently, as deep learning (DL) has undergone remarkable improvements in health care [16], there have
been some efforts toward automating the assessment of digitalized drawing tests such as the pentagon
drawing test (PDT) and the clock drawing test (CDT). In particular, convolutional neural networks (CNN)
that can extract important features automatically from raw data have been widely used and shown
excellent performance. Several automatic scoring systems for PDT were developed using CNN [17–19].
Previous studies on digitalized CDT have shown that CNN could distinguish subjects with cognitive
impairment from cognitively normal (CN) subjects [4, 20, 21].
Meanwhile, for digitalized RCFT, many studies using computer vision technology have been proposed.
For instance, digital tablets and pens have been widely used to generate images and extract distinctive
features from digitalized images. [22] showed the differences between adolescents with attention-de cit
hyperactivity disorder and healthy adolescents by comparing the pixel mean between a template image
and images drawn using a digital tablet. Also, the pen stroke data and spatial information from images
drawn by a digital tablet and pen were also used to distinguish subjects with AD from CN subjects [23].
Furthermore, several DL methods were proposed, recently for the digitalized RCFT. CNN methods using
raw RCFT images have been applied to differentiate individuals with cognitive impairment from those
with CN [24–26]. Those studies have focused on identifying the different patterns between clinical
diagnostic groups and healthy controls.
However, methods for directly predicting RCFT scores based on the 36-point scoring system, which is
widely used in clinical elds, have been very limited. A method to score the RCFT was rstly developed by
segmenting six relevant scoring sections [27]. However, it offered only six of the 18 scoring sections, so
could not be applied to the 36-point scoring system. A DL method for scoring the RCFT was proposed
Page 3/20