Please enable JavaScript.
Coggle requires JavaScript to display documents.
CHAPTER 6 MEASUREMENT OF CONSTRUCTS CHAPTER 7 SCALE RELIABILITY AND…
CHAPTER 6
MEASUREMENT OF CONSTRUCTS
CHAPTER 7
SCALE RELIABILITY AND VALIDITY
Operationalisation refers
to the process of developing indicators or items for measuring these constructs.
Indicators operate at the empirical level, in contrast to constructs, which are
conceptualised at the theoretical level.
The combination of indicators at the empirical level
representing a given construct is called a variable. As noted in the previous chapter, variables may
be independent, dependent, mediating, or moderating, depending on how they are employed in
a research study.
Also each indicator may have several attributes (or levels), with each attribute
representing a different value.
Values of attributes may be quantitative (numeric) or qualitative (non-numeric). Quantitative
data can be analysed using quantitative data analysis techniques, such as regression or structural
equation modelling, while qualitative data requires qualitative data analysis techniques, such as
coding.
Indicators may be reflective or formative.
Reflective indicator is a measure that ‘reflects’
an underlying construct.
Formative indicator is a measure that ‘forms’ or contributes to an underlying construct.
The measure of a construct is consistent or dependable
The extent to which a measure adequately
represents the underlying construct that it is supposed to measure
Conceptualisation is the mental process by which fuzzy and imprecise constructs (concepts) and
their constituent components are defined in concrete and precise terms.
The conceptualisation process is all the more important because of the imprecision, vagueness,
and ambiguity of many social science constructs.
Unidimensional constructs are those that are expected to
have a single underlying dimension. These constructs can be measured using a single measure
or test.
Multidimensional constructs
consist of two or more underlying dimensions.
Conceptualized constructs must be operationalized into measurable indicators.
Validity often called construct validity, refers to the extent to which a measure adequately
represents the underlying construct that it is supposed to measure.
Theoretical assessment of validity focuses on how well the idea
of a theoretical construct is translated into or represented in an operational measure. This type of
validity is called translational validity (or representational validity), and consists of two subtypes:
face and content validity.
Face validity. Face validity refers to whether an indicator seems to be a reasonable measure of
its underlying construct ‘on its face’.
Content validity. Content validity is an assessment of how well a set of scale items matches with
the relevant content domain of the construct that it is trying to measure.
Empirical assessment of validity examines how well a given measure relates to one or more
external criterion, based on empirical observations. This type of validity is called criterion
related validity
Concurrent examines how well one measure relates
to other concrete criterion that is presumed to occur simultaneously.
Predictiveis the degree to which a measure successfully predicts a future outcome that
it is theoretically expected to predict.
Discriminant refers to the degree to
which a measure does not measure (or discriminates from) other constructs that it is not supposed
to measure.
Convergent validity refers to the closeness with which a measure relates to (or converges on)
the construct that it is purported to measure
Reliability the degree to which the measure of a construct is consistent or dependable.
Note that reliability implies consistency but not accuracy.
A measure can be reliable but not valid if it is measuring something very consistently, but is
consistently measuring the wrong construct. Likewise, a measure can be valid but not reliable if
it is measuring the right construct, but not doing so in a consistent manner.
Test-retest reliability. Test-retest reliability is a measure of consistency between two
measurements (tests) of the same construct administered to the same sample at two different
points in time.
Split-half reliability. Split-half reliability is a measure of consistency between two halves of a
construct measure.
Inter-rater reliability. Inter-rater reliability—also called inter-observer reliability—is a
measure of consistency between two or more independent raters (observers) of the same
construct.
Internal consistency reliability. Internal consistency reliability is a measure of consistency
between different items of the same construct.
Hence, it is not adequate
just to measure social science constructs using any scale that we prefer. We also must test these
scales to ensure that: they indeed measure the unobservable construct that we wanted to measure
(i.e., the scales are ‘valid’), and they measure the intended construct consistently and precisely
(i.e., the scales are ‘reliable’). Reliability and validity—jointly called the ‘psychometric properties’
of measurement scales—are the yardsticks against which the adequacy and accuracy of our
measurement procedures are evaluated in scientific research.
Levels of measurement also called rating scales, refer to the values that an
indicator can take (but says nothing about the indicator itself).
Nominal scales, also called categorical scales, measure categorical data.
Ordinal scales are those that measure rank-ordered data, such as the ranking of students in a
class as first, second, third, and so forth, based on their GPA or test scores.
Interval scales are those where the values measured are not only rank-ordered, but are also
equidistant from adjacent attributes.
Ratio scales are those that have all the qualities of nominal, ordinal, and interval scales, and
in addition, also have a ‘true zero’ point (where the value zero implies lack or non-availability
of the underlying construct).
Binary scales.Binary scales are nominal scales consisting of binary items that assume one of two possible values, such as yes or no, true or false, and so on.
Likert scale. Designed by Rensis Likert, this is a very popular rating scale for measuring ordinal
data in social science research.
Semantic differential scale. This is a composite (multi-item) scale where respondents
are asked to indicate their opinions or feelings toward a single statement using different pairs
of adjectives framed as polar opposites.
Guttman scale. Designed by Louis Guttman, this composite scale uses a series of items
arranged in increasing order of intensity of the construct of interest, from least intense to most
intense.
The process of creating the indicators is called scaling. Scaling is
a branch of measurement that involves the construction of measures by associating qualitative
judgments about unobservable constructs with quantitative, measurable metric units.
The outcome of a scaling process is a scale, which is an empirical structure for measuring
items or indicators of a given construct.
Scales can be unidimensional or multidimensional, based on whether the underlying construct
is unidimensional (e.g., weight, wind speed, firm size) or multidimensional (e.g., academic
aptitude, intelligence).
Unidimensional scale measures constructs along a single scale, ranging
from high to low.
The three most popular unidimensional scaling methods
are: Thurstone’s equal-appearing scaling, Likert’s summative scaling, and Guttman’s cumulative
scaling.
Thurstone’s equal-appearing scaling method. Louis Thurstone—one of the earliest and most
famous scaling theorists—published a method of equal-appearing intervals in 1925. 2 This method
starts with a clear conceptual definition of the construct of interest. Based on this definition,
potential scale items are generated to measure this construct.
two additional methods of building unidimensional scales—the method
of successive intervals and the method of paired comparisons—which are both very similar to the
method of equal-appearing intervals, other than the way judges are asked to rate the data.
Likert’s summative scaling method. The Likert method, a unidimensional scaling method
developed by Murphy and Likert (1938), 3 is quite possibly the most popular of the three scaling
approaches described in this chapter.
Guttman’s cumulative scaling method. Designed by Guttman (1950), 4 the cumulative scaling
method is based on Emory Bogardus’ social distance technique, which assumes that people’s
willingness to participate in social relations with other people vary in degrees of intensity, and
measures that intensity using a list of items arranged from ‘least intense’ to ‘most intense’.
Multidimensional scales, on the other hand, employ different items or
tests to measure each dimension of the construct separately, and then combine the scores on
each dimension to create an overall measure of the multidimensional construct.
Well design scale measure of a construct is consistent or dependable.
The items must be adequately representing the underlying construct
Indexes is a composite score derived from aggregating measures of multiple constructs (called
components) using a set of rules and formulas.
It is different from scales in that scales also
aggregate measures, but these measures measure different dimensions or the same dimension of
a single construct. A well-known example of an index is the consumer price index (CPI).
Another example of index is socio-economic status (SES), also called the Duncan socio
economic index (SEI). This index is a combination of three constructs: income, education, and
occupation.
The process of creating an index is similar to that of a scale.
First, conceptualise (define)
the index and its constituent components.
Second, operationalise and measure each component.
Third, create a rule or formula for
calculating the index score.
Consistent Indexes indicators should adequately
represents the underlying construct that it is supposed to measure
indicators must be consistent showing reliability
Typologies summarised measures of two or more constructs to create a set of categories or types.
Unlike scales or indexes, typologies are multidimensional but include
only nominal variables.
An integrated approach to measurement validation
A complete and adequate assessment of validity must include both theoretical and empirical
approaches.
Theory of measurement
Measurement errors can be of two types
Random error
is the error that can be attributed to a set of unknown and uncontrollable external factors that
randomly influence some observations but not others.
Systematic error is an error that is introduced by factors that systematically affect all
observations of a construct across an entire sample in a systematic manner.
Chapter 6
Chapter 7