Perfect colinearity arises from poor sampling when variables mirror each other, preventing the disentanglement of their individual effects. Collecting diverse data with variance resolves this.
◆Main Points
Perfect colinearity can occur due to the specific way a sample is collected.
A size dummy variable equals one if a house has more than two bedrooms.
A country dummy variable equals one if living over 50 miles from a city.
Sampling solely from a wealthy country estate causes size and country to align.
In this flawed sample, large house size perfectly determines the country variable.
Identifying the individual beta coefficients becomes completely impossible under colinearity.
Researchers cannot disentangle the effect of house size from the country effect.
Adding urban poor individuals with small houses keeps variables perfectly aligned.
Both size and country dummies equal zero for the added urban individuals.
The fundamental issue remains a lack of variance between the two variables.
Intelligent sampling requires collecting data with variation across both variables.
House size must not exactly determine whether someone grew up in the country.
✓Takeaways
Poor sampling design can inadvertently create perfect colinearity between distinct variables.
Perfect colinearity makes it impossible to isolate individual variable effects.
Collecting more data is only helpful if it introduces appropriate variance.
Simply adding extreme opposite data points may fail to break colinearity.
Thoughtful sample selection is crucial for valid econometric identification.
Variance across independent variables is essential for statistical disentanglement.
“Quotes
"Perfect colinearity can arise um and it's probably best illustrated by means of an example."
"Our ability to identify beta 1 and beta 2 is going to be impossible."
"We are not going to be able to disentangle the effect of a size of a house from um whether country has any effect."
"This still hasn't enabled us to identify um the effect of a size of house from the effect of being in the country."
"We would hope that we would pick individuals which had some sort of variance in um the size of the house."
"By collecting more data we could essentially get some data whereby the size of the house does not exactly determine um whether they grew up in the country."
⚙Tools
Dummy variables
Beta coefficient estimation
Sample selection methodology
Variance analysis
Population sampling
Data collection expansion
✦Facts
The size dummy uses a threshold of more than two bedrooms.
The country dummy uses a distance threshold of 50 miles to the nearest city.
A wealthy country estate sample resulted in all houses being particularly large.
The urban poor sample featured both country and size variables equaling zero.
Perfect colinearity prevents the identification of beta 1 and beta 2.
The underlying population differs from the poorly extracted sample.
↗References
House size variable (dummy)
Country variable (dummy)
Test score dependent variable
Beta 1 coefficient
Beta 2 coefficient
Wealthy country estate population
→Recommendations
Avoid sampling exclusively from homogeneous subgroups of the population.
Ensure independent variables do not perfectly mirror each other in your dataset.
Introduce variance between variables when expanding your dataset.
Be intelligent about the way in which you collect your samples.
Verify that one variable does not exactly determine another before estimating.
Collect diverse data points rather than just adding extreme opposites.
Summarize any YouTube video — free
Paste a link and get structured notes in seconds. No signup, no card.