Learning Objectives
By the end of this lesson, learners should be able to:
- Define correlation and explain its purpose in business analytics.
- Distinguish between positive, negative and zero correlation.
- Interpret the strength and direction of correlation.
- Explain the limitations of correlation analysis.
- Distinguish correlation from causation.
- Define regression analysis.
- Explain simple and multiple regression.
- Interpret regression coefficients in business contexts.
- Explain the role of the dependent and independent variables.
- Interpret R-squared appropriately.
- Identify potential problems in regression analysis.
- Apply regression concepts to business decision-making.
1. Introduction to Correlation
Correlation measures the degree to which two variables move in relation to one another.
For example, a business analyst might investigate the relationship between:
- Advertising expenditure and sales.
- Customer age and spending.
- Delivery time and customer satisfaction.
- Price and quantity demanded.
- Employee experience and productivity.
Correlation helps identify whether variables tend to move together.
However, correlation by itself does not establish causation.
2. Direction of Correlation
Correlation can generally be:
Positive
Both variables tend to increase together.
Example:
Advertising expenditure increases while sales also tend to increase.
Negative
One variable tends to increase while the other decreases.
Example:
Delivery time increases while customer satisfaction tends to decrease.
Near Zero
There is little evidence of a linear relationship between the variables.
3. Correlation Coefficient
A common measure of linear correlation is the Pearson correlation coefficient, represented by r.
Its value ranges from:
-1 to +1
The sign indicates direction.
The magnitude indicates the strength of the linear relationship.
Interpretation
r = +1
Perfect positive linear relationship.
r = -1
Perfect negative linear relationship.
r = 0
No linear correlation.
Values between these extremes indicate varying degrees of linear association.
4. Interpreting Correlation Strength
Consider the following hypothetical results:
|
Variables |
Correlation |
|
Advertising and sales |
+0.82 |
|
Delivery time and satisfaction |
-0.74 |
|
Age and spending |
+0.18 |
The first relationship indicates a relatively strong positive linear association.
The second indicates a relatively strong negative linear association.
The third indicates a weak positive linear association.
The exact interpretation of “weak” or “strong” should depend on the context rather than relying on rigid universal thresholds.
5. Correlation Does Not Establish Causation
Suppose a company discovers:
Ice cream sales and swimming pool attendance are positively correlated.
It would be inappropriate to conclude that increased ice cream consumption causes people to visit swimming pools.
A third factor may influence both.
In this case:
Hot weather
could increase both ice cream purchases and swimming activity.
This illustrates the importance of considering confounding variables.
6. Spurious Correlation
A spurious correlation occurs when two variables appear related even though the relationship does not represent a meaningful direct relationship.
For example, two unrelated business indicators may move together because both are influenced by a third variable.
Analysts should therefore investigate the business mechanism behind a statistical relationship.
7. Correlation and Business Analysis
Correlation can be useful for:
- Identifying potentially relevant predictors.
- Exploring relationships.
- Supporting hypothesis development.
- Detecting unusual patterns.
- Guiding further analysis.
However, correlation should generally be treated as an exploratory tool rather than definitive evidence of causality.
8. Introduction to Regression Analysis
Regression analysis is a statistical technique used to examine relationships between an outcome variable and one or more explanatory variables.
Unlike correlation, regression is commonly structured around a particular outcome.
For example:
How does advertising expenditure relate to sales?
Here:
- Sales = outcome/dependent variable.
- Advertising expenditure = explanatory/independent variable.
9. Simple Linear Regression
Simple linear regression examines the relationship between:
- One dependent variable.
- One independent variable.
A simplified model can be represented as:
Y = β₀ + β₁X + ε
Where:
- Y = predicted outcome.
- X = predictor.
- β₀ = intercept.
- β₁ = coefficient of X.
- ε = error term.
10. Business Interpretation of a Regression Coefficient
Suppose a regression model estimates:
Sales = 500,000 + 2.4 × Advertising
If advertising expenditure is measured in KSh thousands, the coefficient indicates the estimated change in sales associated with a one-unit increase in advertising expenditure, holding the model structure constant.
The analyst must always understand the units before interpreting a coefficient.
11. The Intercept
The intercept represents the model’s estimated outcome when the predictor equals zero.
In some business situations, the intercept has a meaningful interpretation.
In others, zero may fall outside the realistic range of the predictor.
For example, if a model predicts sales from advertising expenditure, an intercept corresponding to zero advertising may not represent a practically observed situation.
Therefore, the intercept should not automatically be interpreted as a real-world business scenario.
12. Multiple Regression
Multiple regression uses more than one predictor.
For example:
Sales = β₀ + β₁Advertising + β₂Price + β₃Distribution + ε
This allows the analyst to examine how several variables are associated with the outcome simultaneously.
Potential predictors might include:
- Advertising expenditure.
- Product price.
- Distribution coverage.
- Number of sales representatives.
- Seasonality indicators.
13. Why Multiple Regression Is Useful
Business outcomes are rarely determined by one factor.
Suppose sales increase.
Possible contributing factors include:
- Advertising.
- Pricing.
- Distribution.
- Seasonality.
- Competitor activity.
- Economic conditions.
Multiple regression can help account for several variables simultaneously.
However, including more variables does not automatically make a model better.
14. R-Squared
R², or the coefficient of determination, provides an indication of the proportion of variation in the observed outcome that is accounted for by the fitted regression model.
For example:
R² = 0.64
can be described as the model accounting for approximately 64% of the variation in the observed outcome, under the context and assumptions of the model.
It does not mean:
The model is 64% accurate.
This distinction is important.
15. High R-Squared Does Not Prove Causation
A model may have a high R² while still failing to establish a causal relationship.
For example:
Sales and advertising expenditure may have a strong statistical relationship.
But this alone does not prove that advertising caused all of the observed change in sales.
Other factors may be involved.
16. Regression Residuals
A residual represents the difference between an observed outcome and the value predicted by the regression model.
Conceptually:
Residual = Actual value − Predicted value
For example:
Actual sales:
KSh 1,200,000
Predicted sales:
KSh 1,150,000
Residual:
KSh 50,000
Residual analysis can reveal patterns suggesting that the model may not adequately represent the data.
17. Regression Assumptions
Regression analysis relies on assumptions that vary according to the specific model and inference being performed.
Common considerations include:
- Linearity.
- Independence of observations.
- Appropriate treatment of errors.
- Constant error variance in ordinary linear regression.
- Distributional assumptions where required for inference.
Violations can affect interpretation and statistical conclusions.
18. Heteroscedasticity
Heteroscedasticity occurs when the variability of regression errors changes across levels of the predictors or fitted values.
For example:
Predictions for low-value customers may have relatively small errors, while predictions for high-value customers have much larger errors.
This pattern can affect standard errors and statistical inference.
19. Multicollinearity
Multicollinearity occurs when predictors in a multiple regression model are strongly related to one another.
For example:
- Annual income.
- Monthly income.
These variables contain closely related information.
Severe multicollinearity can make individual coefficient estimates difficult to interpret reliably.
20. Regression and Prediction
Regression can be used for prediction.
For example, a retailer could estimate future sales using:
- Historical sales.
- Advertising.
- Price.
- Seasonality.
- Distribution.
However, a regression relationship developed from historical data should not automatically be assumed to remain stable under substantially different future conditions.
21. Extrapolation
Extrapolation occurs when a model is used to make predictions outside the range of data from which the relationship was estimated.
Example:
Historical advertising expenditure:
KSh 100,000–KSh 1,000,000.
An analyst uses the model to estimate sales for:
KSh 10,000,000 advertising expenditure.
That prediction is an extrapolation and may be unreliable if the historical relationship does not continue at that scale.
22. Outliers and Influential Observations
Some observations may have a disproportionate effect on a regression model.
An analyst should investigate unusual observations rather than automatically removing them.
Possible explanations include:
- Data-entry errors.
- Exceptional business events.
- Fraud.
- Genuine extreme customers.
- Structural changes.
23. Correlation Versus Regression
|
Correlation |
Regression |
|
Measures association |
Models an outcome in relation to predictors |
|
Usually symmetric between two variables |
Distinguishes outcome and predictors |
|
Commonly focuses on linear association |
Can be used for estimation and prediction |
|
Does not establish causation |
Does not automatically establish causation |
|
Often useful for exploration |
Often useful for modelling |
Both techniques are valuable, but they serve different analytical purposes.
24. Business Example: Marketing
A company examines:
Advertising expenditure
Sales revenue
Correlation analysis reveals:
r = 0.79
The relationship appears strongly positive.
The company then develops a regression model incorporating:
- Advertising expenditure.
- Price.
- Distribution.
- Seasonality.
The regression model may provide a more useful framework for estimating sales while accounting for multiple variables.
25. Business Example: Customer Satisfaction
A service company examines:
- Waiting time.
- Number of complaints.
- Service satisfaction.
Suppose satisfaction is negatively correlated with waiting time.
This suggests that longer waiting times are associated with lower satisfaction.
However, management should investigate whether:
- Service complexity affects both.
- Certain branches have longer queues and different customer profiles.
- Staffing levels influence both waiting time and service experience.
Correlation identifies a relationship requiring further investigation.
26. Statistical Significance Versus Business Importance
A regression coefficient can be statistically significant without being economically important.
For example:
A model may find that increasing advertising by KSh 1,000 is associated with an average sales increase of KSh 20.
Even if statistically significant, management must determine whether the incremental revenue justifies the advertising cost.
27. Common Regression Mistakes
Analysts should avoid:
- Treating correlation as proof of causation.
- Assuming high R² means a model is automatically good.
- Ignoring outliers.
- Ignoring multicollinearity.
- Extrapolating far beyond the observed data.
- Interpreting coefficients without considering units.
- Including variables without considering their business meaning.
- Assuming historical relationships will remain unchanged.
28. Best Practices
Business analysts should:
- Define the business question clearly.
- Understand the data before modelling.
- Examine relationships visually and statistically.
- Investigate unusual observations.
- Consider relevant variables.
- Check appropriate regression assumptions.
- Interpret coefficients in business terms.
- Avoid causal claims without appropriate evidence.
- Evaluate predictive performance using appropriate validation.
- Communicate limitations clearly.
Lesson Summary
Correlation provides a measure of the direction and strength of a linear relationship between variables.
Regression goes further by modelling an outcome in relation to one or more predictors.
Important concepts include:
- Positive and negative correlation.
- Correlation coefficient.
- Spurious correlation.
- Causation versus association.
- Simple regression.
- Multiple regression.
- Regression coefficients.
- Intercepts.
- R².
- Residuals.
- Multicollinearity.
- Heteroscedasticity.
- Extrapolation.
- Outliers.
The central principle is:
A statistical relationship can provide evidence of association without automatically establishing causation.