Heteroskedasticity caused by data aggregation (advanced topic)
AI-extracted key points, takeaways & quotes
Aggregating individual data into groups can improve model fit but sacrifices estimator efficiency and introduces heteroscedasticity because error variances differ across unequally sized groups.
◆Main Points
Aggregating data means grouping individual observations, such as students into schools.
Individual data contains many idiosyncratic error sources that lower model fit.
Grouping data averages out individual variations with a mean of zero.
Aggregated data often appears to fit a model significantly better than individual data.
There are no free lunches in econometrics; aggregation introduces hidden statistical problems.
Moving from individual to group level data discards important individual sources of variation.
Least squares estimators of beta lose efficiency when individual variation is thrown away.
Aggregated estimators may remain unbiased but are no longer efficient.
Individual level data generally yields estimates closer to the true population parameter.
Heteroscedasticity in aggregated data arises directly from the behavior of the error term.
A group's average error is calculated by weighting the linear sum of individual errors.
Unequal group sizes cause the variance of average errors to differ across groups.
✓Takeaways
Better model fit in aggregated data does not guarantee a better statistical model.
Data aggregation inherently destroys valuable individual-level variance.
Efficiency in least squares estimators is sacrificed when data is grouped.
Heteroscedasticity is fundamentally about the unequal size or variance of error terms.
The variance of a group's error depends inversely on the number of individuals in it.
Using differently sized groups guarantees that error variances will be unequal.
“Quotes
"There are no free lunches in econometrics."
"By throwing away this sort of individual source of variation, we are sacrificing their efficiency."
"Whenever we are talking about heteroscedasticity, we're talking about heteroscedasticity in the size of the error term."
"The average error of individuals in a school is made up of a linear sum of the errors of all individuals."
"The variance of our errors is not equal to some constant sigma squared. It varies by group."
"If the groups are made of different numbers of individuals, that means that the variance of errors in both of these groups is going to be different."
⚙Tools
Least squares estimators
Data aggregation (grouping)
Linear regression modeling
Error variance analysis
Weighted sum calculations
Population parameter estimation
✦Facts
Aggregation reduces 10,000 individual observations down to 500 group observations.
A group's average error equals the sum of its individual errors divided by the group size.
Heteroscedasticity literally means the variance of errors is not equal to a constant.
Aggregated models can still produce unbiased estimates despite losing efficiency.
The variance of a group's error changes based on the number of individuals within it.
Individual-level data contains idiosyncratic errors that average out to zero when grouped.
↗References
Test scores as a dependent variable
Parental income as an independent variable
School-level grouping
Population parameter beta
Constant variance sigma squared
Individual sources of variation
→Recommendations
Prefer individual-level data over aggregated data to maintain estimator efficiency.
Do not be misled by the apparently better fit of aggregated regression models.
Always check for heteroscedasticity when analyzing grouped or aggregated data.
Account for varying group sizes when estimating models with aggregated observations.
Retain individual variation in your datasets rather than discarding it through grouping.
Remember that unbiased estimates can still be statistically inefficient and suboptimal.
Summarize any YouTube video — free
Paste a link and get structured notes in seconds. No signup, no card.