Reliability assessment
Estimate score consistency under specified replications
Reliability concerns consistency of scores over specified replications in a population. Test-retest evidence addresses temporal stability; inter-rater evidence addresses variation between judges; parallel-forms evidence concerns interchangeable forms; internal consistency concerns components within a scale. These differ in their error sources. For ratings, the intraclass correlation model must distinguish agreement from consistency and single from averaged measurements.
Use when measurement error could affect interpretation and you can design repeated administrations, independent ratings or alternative forms that reflect the intended use of the score.
Strengths
- Quantifies an important source of uncertainty in scores
- Helps decide whether averaging ratings or improving procedures is useful
Limitations
- Estimates depend on population variability and replication conditions
- Temporal change can be confused with measurement error
Know the boundary
Reliability is not accuracy against truth and cannot establish validity on its own.