Reliability determines how consistently a measurement of skill or knowledge yields similar results under varying conditions. If a measure has high reliability, it yields consistent results. There are four principal ways to estimate the reliability of a measure:
Inter-observer: Is determined by the extent to which different observers or evaluators examine the same presentation, demonstration, project, paper, or other performance and agree on the overall rating on one or more dimensions.
Test-retest: Is determined by the extent to which the same test items or kind of performance evaluated at two different times yields similar results.
Parallel-forms: Is determined by examining the extent to which two different measurements of knowledge or skill yield comparable results.
Split-half reliability: Is determined by comparing half of a set of test items with the other half and determining the extent to which they yield similar results.