Please enable JavaScript.
Coggle requires JavaScript to display documents.
Hypothesis Testing - Coggle Diagram
Hypothesis Testing
-
Influences
The variability of scores - Increased variability means that the sample data are no longer sufficient to conclude that the treatment has a significant effect. If other factors are held constant, the larger the variability, the lower the likelihood of finding a significant treatment effect.
A significant (or statistically significant) result means that the null hypothesis has been rejected. It is very unlikely to occur when the null hypothesis is true. It is expressed using z-scores (z) and probabilities (p).
The number of scores in a sample - A decreased number of scores in the sample produces a larger standard error and a smaller value (closer to zero) for the z-score. If all other factors are held constant, a larger sample is more likely to result in a significant treatment effect.
Errors
Type I
Happens when we reject a true null hypothesis. Extreme values in a sample can give the impression that a treatment had an effect when it actually didn’t. Whenever a researcher rejects the null hypothesis, there is a risk of a Type I error.
The alpha level for a hypothesis test is the probability that the test will lead to a Type I error. By selecting a small alpha level, the researcher can minimize the probability of a Type I error.
Alpha Levels
-
More conservative Alpha Levels can minimize the damage of a Type I error, which creates the need for more evidence; this trade-off is controlled by the boundaries of the critical region. Sample data must be in the critical region.
Alpha levels of .05, .01, and .001 are considered reasonably good values.
As the alpha level gets smaller, this distance between the sample mean and the population mean gets larger.
Type II
Failing to reject a false null hypothesis. This occurs when a treatment effect really exists, but the hypothesis test fails to detect it. A Type II error can be attributed to many factors and we can’t determine the probability of making it (but still, that probably is represented by B)
This most often happens when the effect of the treatment is relatively small.
-
Statistical power
The power of a statistical test is the probability that the test will correctly reject a false null hypothesis.
Calculating power
the probability that the test will correctly reject the null hypothesis. The value chosen for the alpha level can influence a hypothesis’power. We can change an assumption of the power analysis to see how the degree of power is influenced. When the effect size is small, the statistical power will be low—provided that other factors are held constant
.
Power and effect size
As the effect size increases, the distribution of sample means for the alternative distribution moves even farther from the null distribution. The chances of rejecting the null hypothesis also increases. power analysis will sometimes accompany Cohen’s d and similar measures of effect size.
Factors affecting power
Reducing alpha levels reduces the test’s power; A one-tailed test has greater power than a two-tailed test.
Power and sample size
Power is directly related to sample size. A larger sample produces greater power for a hypothesis test.
-
As a ratio, formed through a z-score
Formula for z-score as a ratio in hypothesis testing: s\Sample mean minus hypothesized sample mean divided by the sample error between M and ų;the z-score is the actual difference between M and ų divided by the expected difference between M and ų with no treatment effect. Large value = large discrepancy, which must be caused by treatment effec
Measure of effect size is intended to provide a measurement of the absolute magnitude of a treatment effect, independent of the size of the sample(s) being used.
Cohen's d measures the distance between two means and is typically reported as a positive number even when the formula produces a negative value.
Formula for measuring effect size (Cohen’s d): The mean of difference (difference between mean of treatment and mean of no treatment- the null hypothesis) divided by the standard deviation. We must use the mean for the treated sample in place of the population mean, since that is unknown.
If the z-score is large enough to be in the critical region, we reject the null hypothesis and conclude that there is a significant treatment effect. Otherwise, we fail to reject and conclude that the treatment does not have a significant effect
The most obvious factor influencing the size of the z-score is the difference between the sample mean and the hypothesized population mean from