Please enable JavaScript.
Coggle requires JavaScript to display documents.
Simple linear regression 2 - Coggle Diagram
Simple linear regression 2
Prediction error and variance partitioning
Residuals and prediction error
A residual is the difference between an observed outcome and its predicted value:
Residual = Y − Ŷ
A positive residual means the observed outcome is higher than predicted. A negative residual means the observed
outcome is lower than predicted. Smaller residuals indicate that predictions are closer to the observed values (Usually when correlation is strong)
H
The mean-only model, often called the null model, is the simplest possible baseline in statistics. It predicts the exact same value—the average (mean) of the outcome variable—for every single observation, completely ignoring any other information or predictor variables.
Here is a breakdown of how it works and why it matters in simple regression:
💡 The Core Concept
Imagine you are trying to guess the weight of 100 random people, but you know absolutely nothing about them. Your best, safest guess for every single person is the average weight of the entire group.
The Model: \(Y = \bar{Y}\) (where \(\={Y}\) is the sample mean).
The Logic: If you don't use a predictor (like height), the mean is the mathematically optimal single number that minimizes the total squared guessing errors.
H
⚖️ Why We Use It: The Ultimate Baseline
The primary job of a regression model is to prove it is useful. The mean-only model serves as the benchmark or starting line for evaluation.
When you build a simple regression model (using a predictor like height to guess weight), you are asking: "Does adding this predictor actually help us make better guesses than just using the flat average?"
To see if your regression model is successful, you compare its errors against the null model's errors:
Total Sum of Squares (SST): This measures the total error or variation when using only the mean-only model.
Residual Sum of Squares (SSE): This measures the remaining error after you introduce your predictor variable.
\(R^{2}\) (Coefficient of Determination): This calculates the percentage of error you successfully eliminated by moving from the mean-only model to your regression model. If your regression model doesn't beat the mean-only model, your predictor is practically useless.
An R² of .335 means that the model accounts for 33.5% of the variance in the outcome. The remaining 66.5% is not
accounted for by the model. R² describes the amount of variance accounted for; it does not indicate whether the
model is statistically significant and it does not establish causation.
Adjusted R² takes the number of predictors and the sample size into account. It provides a more cautious
estimate and is particularly useful when comparing models containing different numbers of predictors
R² is the coefficient of determination. It represents the proportion of outcome variance accounted for by the
regression model:
R² = SSM ÷ SST
Model fit tells you how well your statistical model replicates the actual real-world data you collected. While statistical significance answers the question, "Is there a real relationship here, or is this just random luck?", model fit answers, "How accurately can this model predict or explain the outcome?"In simple linear regression, you can evaluate model fit using four key indicators
1. R² (Coefficient of Determination)
What it is: The percentage of the variation in the dependent variable (\(Y\)) that can be explained by the independent variable (\(X\)).
How to interpret it: It ranges from 0 to 1 (or 0% to 100%). If your R² is 0.75, it means your model explains 75% of the variance in your data, leaving 25% unexplained.
2. Adjusted R²
What it is: A modified version of R² that accounts for the number of predictors in a model.
How to interpret it: While more relevant in multiple regression, it is important because standard R² artificially increases every time you add a new variable—even if that variable is useless. Adjusted R² penalizes unnecessary complexity, ensuring your model fit is genuine.
3. Residual Size
What it is: A "residual" is the error margin—the specific distance between an actual observed data point and the prediction line created by your model.
How to interpret it: Smaller residuals mean the model's guesses are very close to reality. Standard metrics like the Residual Standard Error (RSE) quantify the average size of these deviations.
4. Observed vs. Predicted Values
What it is: A visual or mathematical comparison of what actually happened versus what your regression line predicted would happen.
How to interpret it: If you plot your observed values against your predicted values on a graph, a perfect model would form a straight, tight diagonal line. Deviations or weird curves in this plot signal that your model is missing key patterns.
H
The F-test evaluates whether a regression model as a whole predicts an outcome significantly better than a basic model that only uses the mean. [1]
How the F-Test Works
Model Comparison: It checks if your predictor variables add useful predictive power compared to a baseline (mean-only or intercept-only) model.
The F-Statistic: Calculated as \(F = \text{MS}_{\text{model}} \div \text{MS}_{\text{residual}}\) (model mean square divided by residual mean square). A higher \(F\) value means your model explains more variance than the left-over random error.
The P-Value: Tells you if the improvement in prediction is statistically significant or just due to random chance
What the F-Test Does Not Do
It does not show effect size: It does not tell you the exact amount of variance accounted for (like \(R^{2}\) does).
It does not prove causation: A strong predictive relationship does not mean the predictors caused the outcome.
A confidence interval (CI) is a range of values that likely contains the true population parameter based on your sample data. It tells you not just an estimate, but how precise or uncertain that estimate is.
Example-(B = 0.096), 95% CI [(0.077, 0.115)].
Plausible Values (The Range)- The point estimate (\(B = 0.096\)) is the exact effect found in your specific sample. However, because samples vary, the true population value could be slightly higher or lower. The 95% CI tells us that values between 0.077 and 0.115 are highly plausible estimates for the true population coefficient.
2. Precision and Uncertainty (The Width)- The width of the interval directly communicates the reliability of your data:
Narrow interval: Indicates high precision. You have a very good idea of what the true population value is.
Wide interval: Indicates high uncertainty. The data is noisy, or the sample size is too small to pinpoint the true value accurately.
In regression, a coefficient of \(0\) means the predictor variable has absolutely no effect on the outcome.
3. Statistical Significance (The Zero Rule)- If a confidence interval does not contain zero (like your interval of \(0.077\) to \(0.115\)), it means zero is an implausible value. Therefore, the predictor contributes significantly to the model (\(p < 0.05\)).
If the interval does contain zero (e.g., \([-0.02, 0.11]\)), you cannot rule out the possibility that the true effect is zero. Therefore, the predictor is not statistically significant.
Direction of the Relationship
Because every single value inside your interval [(0.077, 0.115)] is a positive number, you can confidently conclude that the estimated relationship is positive. As your predictor variable increases, the outcome variable increases. If all values were negative, the relationship would be strictly negative.
Before interpreting the findings, researchers should evaluate whether a straight line adequately represents the
relationship and whether unusual observations affect the model. A statistically significant model is not
automatically a trustworthy model.
• Linearity: Does a straight line adequately represent the relationship between X and Y?
• Outliers: Are any observations unusually extreme?
• Influential cases: Would removing a particular observation meaningfully change the fitted model?
• Residuals: Are prediction errors sufficiently well behaved for the model to be interpreted?
Outliers and influential cases are related but distinct. An outlier is unusual, whereas an influential case has a
substantial effect on the estimated model. An observation can be influential without being an obvious univariate
outlier. Unusual observations should be investigated rather than removed automatically