Back to Glossary

Multicollinearity

Multicollinearity is a statistical problem where two or more explanatory variables in a model move together so closely that the model cannot tell their effects apart. It inflates uncertainty and makes individual coefficients unreliable. It is one of the most common hidden flaws in business data analysis.

From the Money Master HQ dictionary, founded by Shihan Sheriff (FCMA, VP of Finance at Nomod, CFO at Esanjo Ventures). How these definitions are written.

What it means

A regression tries to assign credit: how much of the outcome came from each input? Multicollinearity is what happens when two inputs always travel together, leaving the model unable to say which one did the work.

The classic illustration is absurd and memorable. Predict foot size from both left-shoe size and right-shoe size, and the model falls apart, because the two measures carry the same information and the credit can be split arbitrarily.

Real business models hit the same wall. Advertising spend and promotional discounts that always run together, or staffing and opening hours that rise in lockstep, leave the analysis unable to separate their effects on sales.

The diagnosis has standard tools. Statisticians inspect correlation among predictors and compute variance inflation factors, and university statistics courses, such as Pennsylvania State University's regression materials, teach the detection and the remedies as core technique.

The fixes are practical: drop one of the entangled variables, combine them into a single measure, gather data where they move separately, or accept that the model can predict well even while individual effects stay blurry. For a business owner, multicollinearity is a warning about dashboard analytics.

A report claiming to isolate what drives sales can be quietly confounded when the supposed drivers always move together, and the remedy is usually an experiment, not a bigger spreadsheet. The problem worsens as models grow.

Every added variable is another chance to import a twin of one already there, and kitchen-sink regressions built from whole data warehouses are almost guaranteed to suffer it. Machine learning has not abolished the issue.

Flexible models absorb correlated inputs gracefully for prediction, but the moment someone asks which feature drove the decision, the same old attribution problem returns wearing new clothes. Practitioners keep a simple habit: before believing any driver analysis, ask how the supposed drivers relate to each other, because the interesting question is rarely whether they predict, but whether they can be told apart.

In practice

Real-world examples.

1

Example

A retailer models sales against TV spend and social spend, which always ran in the same weeks. The coefficients swing wildly between runs, and the analyst concludes the two channels cannot be separated from history alone. She recommends a short test in which the spends are varied independently.

2

Example

A bank's risk model uses income and credit score, which correlate strongly. Dropping one stabilises the model, and validation confirms nothing was lost. The model team documents the correlation and the reason for the exclusion.

3

Example

A restaurant chain runs a deliberate test: half its branches discount while half advertise, decoupling the variables for eight weeks. The clean data finally prices each effect, ending a year of argument. Management uses the results to set next year's promotional budget.

Formula

Calculation

Variance inflation factor = 1 / (1 - R-squared) from regressing one predictor on the others. A VIF of 1 means independence; above 5 or 10 signals trouble. Two variables correlated at 0.9 have an R-squared of 0.9 x 0.9 = 0.81, so the VIF is 1 / (1 - 0.81) = 1 / 0.19, which is about 5.3. The standard error of a coefficient is inflated by the square root of the VIF. The square root of 5.3 is about 2.3, so the standard errors of both coefficients are more than twice as wide as they would be with uncorrelated predictors. A coefficient that looked clearly different from zero can therefore become statistically indistinguishable from zero.

Case study

Seen in the real world.

In this illustrative fictional case, Tarik, analytics lead at a subscription software firm, presents a churn model assigning precise effects to onboarding calls, feature usage and support tickets. A peer reviewer spots that engaged customers simply do more of everything, and the predictors' VIFs exceed 8. The team rebuilds the model with a single engagement score for prediction, and separately runs controlled trials to isolate which lever truly causes retention. The board gets two honest answers instead of one false one: what predicts churn, and what changes it. Tarik's slide for every future model review is one line: correlated inputs make confident nonsense.

Watch out

Common mistakes.

  • Trusting individual coefficients when predictors move together, when multicollinearity makes their estimates unstable and their signs arbitrary.
  • Assuming the model is useless, when multicollinearity blurs attribution but often leaves overall prediction perfectly serviceable.
  • Adding more correlated variables to fix it, when the remedies are dropping, combining, or collecting data where the drivers actually move apart.

Questions

People also ask.

What causes multicollinearity?

Explanatory variables that carry the same information: measures that rise and fall together by nature or by habit, such as two marketing channels always run in the same weeks.

How do you detect it?

Check correlations among predictors and compute variance inflation factors, the standard diagnostic taught in university regression courses. VIF values above 5 to 10 flag a problem worth fixing.

Does multicollinearity ruin a model?

It ruins interpretation, not necessarily prediction. The model may forecast well while being unable to say which variable deserves the credit, so the fix depends on whether you need predictions or levers to pull.

Was this explanation helpful?

From the founder's library

Accounting Fundamentals: A Non-Finance Manager's Guide to Finance and Accounting, by Shihan Sheriff

Take it further with the book.

Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.

US$2.24US$2.99

25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.

View the book and save 25%
Last updated · October 8, 2026
Browse all terms →

Disclaimer

The information provided in this finance dictionary is for educational and informational purposes only. It should not be construed as financial, investment, legal, or tax advice. Always consult with a qualified professional before making any financial decisions. Money Master HQ makes no representations or warranties about the accuracy, completeness, or suitability of this information. Use of this content is at your own risk.