DOI 10.1007/s11136-011-0076-4
Development of an item bank and computer adaptive test for role
functioning
Milena D. Anatchkova • Matthias Rose •
John E. Ware Jr. • Jakob B. Bjorner
Accepted: 21 November 2011
Springer Science+Business Media B.V. 2011
Abstract and occupational). Slopes in the bank ranged between .93
Objectives Role functioning (RF) is a key component of and 4.37; the mean threshold range was -1.09 to -2.25. Item
health and well-being and an important outcome in health bank-based scores were significantly different for partici-
research. The aim of this study was to develop an item pants with and without chronic conditions and with different
bank to measure impact of health on role functioning. levels of self-reported general health.
Methods A set of different instruments including 75 Conclusions An item bank assessing health impact on RF
newly developed items asking about the impact of health across three content areas has been successfully developed.
on role functioning was completed by 2,500 participants. The bank can be used for development of short forms or
Established item response theory methods were used to computerized adaptive tests to be applied in the assessment
develop an item bank based on the generalized partial of role functioning as one of the common denominators
credit model. Comparison of group mean bank scores of across applications of generic health assessment.
participants with different self-reported general health
status and chronic conditions was used to test the external Keywords Role function IRT Computer adaptive test
validity of the bank. CAT
Results After excluding items that did not meet established
requirements, the final item bank consisted of a total of 64
items covering three areas of role functioning (family, social, Introduction
Social well-being is one of the main aspects of health
M. D. Anatchkova (&) M. Rose J. E. Ware Jr. recognized by the World Health Organization in its defi-
Department of Quantitative Health Sciences, Medical School,
nition of health as not only the absence of disease, but also
University of Massachusetts, 55 Lake Avenue North, Worcester,
MA, USA the presence of physical, mental, and social well-being [1].
e-mail: Role functioning is a key component of social well-being
with impaired role functions and role disability recognized
M. Rose
as major sources of indirect cost of illness and disease
Department of Psychosomatic Medicine, University Clinic
Hamburg-Eppendorf and Schön Klinik Hamburg-Eilbek, burden [2], making role functioning an important construct
Hamburg, Germany to assess accurately.
In recent work, we have focused on the development of
J. E. Ware Jr.
a generic item bank assessing the impact of health on role
John Ware Research Group, Worcester, MA, USA
functioning using a previously described multistage pro-
J. B. Bjorner cess [3, 4]. Our theoretical conceptualization of the role
National Institute of Occupational Health, functioning construct was inspired by the biopsychosocial
Copenhagen, Denmark
model of health and disability and the International
J. B. Bjorner Classification of Functioning, Disability and Health ICF
i3 QualityMetric Inc, Lincoln, RI, USA [1]. We defined role functioning as involvement in life
123
, Qual Life Res
situations related to family life, partner relationship, consent form and received an incentive for their partici-
household chores, work for pay, studies, social life pation in the study.
(including interactions with friends), leisure time activities,
community involvement (including volunteer work), and Instruments
everyday living activities [5]. The qualitative stages of the
item development resulted in the formulation of 75 items The bank included 75 newly developed items assessing
assessing the health impact of role functioning across 3 role role functioning and 4 items from the Role Physical Scale
areas, e.g. family life, occupational life, and social life [5] of the SF-36 Health Survey, Version 2 [13]. Review of
(see Table 1 for item text). We have reported the factor existing measures and focus groups were used in the item
structure of the bank resulting in a model with one overall development process [5]. All items had a 4-week recall
general factor covering three content areas (social, family, period and were presented consecutively in a one item per
and occupation) [6]. Based on this work, we concluded that screen format. Standard procedures were used to ensure
the item response theory (IRT) requirements of sufficient and evaluate unidimensionality and local independence of
unidimensionality and local independence (i.e. no residual the items as described earlier [6].
correlations [ 0.02 in the confirmatory factor model) of
the items have been fulfilled and the application of IRT for Statistical analyses
further analyses is appropriate [6].
Item response theory models allow the selection of the Descriptive statistics were computed for the demographic
most appropriate items for each respondent through com- characteristics of the sample and for each individual item
puterized adaptive testing (CAT) [7–10] of role function- of the item bank. Item characteristic curves (ICC) were
ing. This is a promising new strategy in the development of evaluated using the TestGraf program [14], inspecting for
improved health outcomes measures [11, 12]. The flexi- monotonicity of response curves. The ICC represents the
bility of CAT is of particular interest for constructs where probability, for each item, of selecting a response choice
one underlying dimension can be identified, but several category for each level of the latent trait (e.g. impact of
distinctive areas are also of interest as is the case for the health on role functioning). Good items should have
role functioning item bank. Calibration of all items on the response choice categories with unequivocal and unique
metric of the general underlying dimension (e.g. the impact relationships to the latent trait organized in rank order (e.g.
of health on role functioning) allows the computation of the response curve for response ‘‘most of the time’’ should
one overall impact score, while at the same time, sampling be between the curves of responses ‘‘some of the time’’ and
items from all relevant content areas (e.g. impact on social, ‘‘all of the time’’). There should be an interval on the latent
occupational, and family functioning) using item selection trait where a specific response choice category is most
algorithms that account for content. likely to be selected. Visual inspection of these ICC graphs
This paper describes the results of the next steps in item provides insight into whether an item’s response choice
bank development: the calibration of all items on a com- categories overlap and require re-scaling (or collapsing
mon metric using IRT methods and the evaluation of the across response choice categories) to achieve distinct cat-
bank’s properties and validity. egories [3].
To evaluate whether items function in the same manner
for participants regardless of group/respondent character-
Methods istics, we conducted a test of differential item functioning
(DIF) by demographic variables (gender, age, and chronic
Participants condition). Multiple approaches to the evaluation of DIF
have been proposed and reviewed in the literature [15–17].
A total of 2,500 participants were recruited for the As in previous studies, we used logistic regression—one of
study via the Internet by a panel company (www. the established parametric approaches to DIF analyses [4,
YouGovPolimetrix.com). The sample was stratified by 11]. This approach tests the association between each item
age group with equal gender representations and race and and subgroup membership, while controlling for the simple
ethnicity quotas representative of the US population. sum score of all the items that are found to be without DIF
Approximately half of the sample (n = 1,445) included in initial analyses. DIF is indicated by a significant asso-
participants with selected self-reported chronic conditions ciation between subgroup membership and item score
(asthma n = 416, heart disease n = 333, diabetes n = 367, when controlling for the overall sum score. This approach
auto-immune n = 329), and the other half was recruited allows for the simultaneous testing of uniform DIF (tested
from the general population, but excluded persons with through a logistic regression main effect and similar to
these conditions (n = 1,055). All participants completed a differences in IRT threshold parameters) and non-uniform
123