NCE Assessment
1. Used to estimate the impact that shortening or lengthening a test will have on its reliability:
Spearman-Brown Formula
2. Measure of internal consistency that compares item responses with eachother and the
total test score: Interitem Consistency
3. Used to calculate interitem consistency when items are dichotomous(yes/no, true/false):
Kuder-Richardson Formula
4. Used to calculate interitem consistency when items have multipoint re-sponses (multiple
choice): Cronbach's Coefficient Alpha
5. Degree to which two people give consistent ratings when viewing the samebehavior: Inter-
Scorer Reliability
6. Correlation that expresses a test's reliability; the closer to 1.0 the better;also called
Pearson r; range is -1 to +1: Reliability Coefficient
7. Standard deviation of a persons' repeated test scores; inversely related to reliability (ex.
reliability = 1.0, SEM = 0); helps you figure out what will probably happen if the subject takes the
same test again: Standard Error of Measurement
8. Can test scores be reliable but not valid?: Yes
9. Can test scores be valid but not reliable?: No
10. Statistically examining people's item responses to assess the quality ofthe test: Item
Analysis
11. % of test takers who answer an item correctly; it's a p value between 0 and1 with a high
value meaning an easier item; .5 is ideal: Item Difficulty
12. Degree to which a test item differentiates test-takers; ex. does item on depression
inventory get a different answer from depressed people versusnon-depressed people?: Item
Discrimination
13. Theory that says score = true score + error: Classical Test Theory
14. Theory that uses mathematical models to detect item bias or equate scoresfrom two different
tests: Item Response Theory
15. Scale that classifies or label, has no zero point, and doesn't indicate order;ex. gender:
Nominal
16. Scale that shows rank order but intervals between numbers aren't equal;ex. places in a
horse race: Ordinal
17. Scale that has numbers at equal distances but no absolute zero; ex.Fahrenheit: Interval
18. An interval scale with a true zero point; ex. height, weight: Ratio
19. Scale that asks test takers to place a mark between two dichotomousadjectives; also called
, self-anchored scale: Semantic Differential
20. Scale that measure multiple dimensions of an attitude by asking test takersto agree/disagree
with a series of statements: Thurstone Scale
21. Scale that measures the intensity of a variable in progressive order; ex.would you permit
gay students to live on campus? would you have a gay roommate?: Guttman Scale
22. Converted raw score that gives meaning by comparing to a norm group: -
Derived Score
23. Characteristic of normal curve where tail approaches horizontal axis with-out ever
touching it: Asymptotic
24. Assessment that compares a person's score to a pre-determined standard;ex. NCE, drivers
license test: Criterion Referenced Assessment
25. Assessment that compares a person's score to a previous test score; ex.PE class, computer
game: Ipsative Assessment
26. Developed first intelligence theory: Galton
27. Developed first modern intelligence test and coined term IQ: Binet
28. Revised Binet's IQ test into the Stanford-Binet: Terman
29. Mental age divided by chronological age x 100: Intelligence Quotient
30. Test that has no time limit and includes difficult items that few test takerscan answer:
Power Test
31. Test with a time limit; usually have easy items but too many to answer intime limit:
Speed Test
32. Test that allows an individual's score to be compared to a norm group: Stan-dardized Test
33. Best source of information about commercially available assessments; pro-vides reviews
of tests; has companion called Tests in Print: Mental Measure- ments Yearbook
34. Overview of assessments for the layperson: Test Critiques
35. Developed by Robert Yerkes to screen cognitive ability of military recruits;group
intelligence test: Army Alpha and Beta
36. Term that refers to whether test measures what it's supposed to measure;depends on test
purpose and target population: Validity
37. A depression inventory that has ?s on all the aspects of depression (physical, emotional,
cognitive) has what kind of validity?: Content Validity
38. Type of validity that shows how effective an instrument is at predicting anindividual's
performance; two kinds are concurrent and predictive: Criterion Validity
1. Used to estimate the impact that shortening or lengthening a test will have on its reliability:
Spearman-Brown Formula
2. Measure of internal consistency that compares item responses with eachother and the
total test score: Interitem Consistency
3. Used to calculate interitem consistency when items are dichotomous(yes/no, true/false):
Kuder-Richardson Formula
4. Used to calculate interitem consistency when items have multipoint re-sponses (multiple
choice): Cronbach's Coefficient Alpha
5. Degree to which two people give consistent ratings when viewing the samebehavior: Inter-
Scorer Reliability
6. Correlation that expresses a test's reliability; the closer to 1.0 the better;also called
Pearson r; range is -1 to +1: Reliability Coefficient
7. Standard deviation of a persons' repeated test scores; inversely related to reliability (ex.
reliability = 1.0, SEM = 0); helps you figure out what will probably happen if the subject takes the
same test again: Standard Error of Measurement
8. Can test scores be reliable but not valid?: Yes
9. Can test scores be valid but not reliable?: No
10. Statistically examining people's item responses to assess the quality ofthe test: Item
Analysis
11. % of test takers who answer an item correctly; it's a p value between 0 and1 with a high
value meaning an easier item; .5 is ideal: Item Difficulty
12. Degree to which a test item differentiates test-takers; ex. does item on depression
inventory get a different answer from depressed people versusnon-depressed people?: Item
Discrimination
13. Theory that says score = true score + error: Classical Test Theory
14. Theory that uses mathematical models to detect item bias or equate scoresfrom two different
tests: Item Response Theory
15. Scale that classifies or label, has no zero point, and doesn't indicate order;ex. gender:
Nominal
16. Scale that shows rank order but intervals between numbers aren't equal;ex. places in a
horse race: Ordinal
17. Scale that has numbers at equal distances but no absolute zero; ex.Fahrenheit: Interval
18. An interval scale with a true zero point; ex. height, weight: Ratio
19. Scale that asks test takers to place a mark between two dichotomousadjectives; also called
, self-anchored scale: Semantic Differential
20. Scale that measure multiple dimensions of an attitude by asking test takersto agree/disagree
with a series of statements: Thurstone Scale
21. Scale that measures the intensity of a variable in progressive order; ex.would you permit
gay students to live on campus? would you have a gay roommate?: Guttman Scale
22. Converted raw score that gives meaning by comparing to a norm group: -
Derived Score
23. Characteristic of normal curve where tail approaches horizontal axis with-out ever
touching it: Asymptotic
24. Assessment that compares a person's score to a pre-determined standard;ex. NCE, drivers
license test: Criterion Referenced Assessment
25. Assessment that compares a person's score to a previous test score; ex.PE class, computer
game: Ipsative Assessment
26. Developed first intelligence theory: Galton
27. Developed first modern intelligence test and coined term IQ: Binet
28. Revised Binet's IQ test into the Stanford-Binet: Terman
29. Mental age divided by chronological age x 100: Intelligence Quotient
30. Test that has no time limit and includes difficult items that few test takerscan answer:
Power Test
31. Test with a time limit; usually have easy items but too many to answer intime limit:
Speed Test
32. Test that allows an individual's score to be compared to a norm group: Stan-dardized Test
33. Best source of information about commercially available assessments; pro-vides reviews
of tests; has companion called Tests in Print: Mental Measure- ments Yearbook
34. Overview of assessments for the layperson: Test Critiques
35. Developed by Robert Yerkes to screen cognitive ability of military recruits;group
intelligence test: Army Alpha and Beta
36. Term that refers to whether test measures what it's supposed to measure;depends on test
purpose and target population: Validity
37. A depression inventory that has ?s on all the aspects of depression (physical, emotional,
cognitive) has what kind of validity?: Content Validity
38. Type of validity that shows how effective an instrument is at predicting anindividual's
performance; two kinds are concurrent and predictive: Criterion Validity