What it means
For two categorical variables, a test of independence asks whether their joint counts depart from what independence would predict. An insurer might compare policy type with whether a claim occurred.
The test indicates evidence of association, not its direction or cause. Start with a clearly stated null hypothesis.
For goodness of fit, specify expected proportions before looking at the results; for independence, the expected count in a table cell is its row total times column total divided by the overall total. The statistic is a sum over categories or cells, where squaring prevents positive and negative deviations from cancelling and dividing by the expected count scales a difference relative to the frequency predicted for that category.
OpenStax gives different degrees of freedom for the two tests. A basic goodness-of-fit test with k fixed categories uses k minus one, and a table with r rows and c columns uses (r minus one) times (c minus one).
Estimating model parameters from the same data can change the goodness-of-fit degrees of freedom. The p-value measures how surprising a statistic at least this large would be if the null model and assumptions held.
A small p-value is evidence against the null, not the probability that the null itself is false, and a chosen significance threshold should be set before inspecting results. These tests use counts of observations in mutually exclusive categories, not averages or percentages entered as if each were a raw observation, and one person or loan should not be counted as independent several times without an appropriate design.
The independence assumption generally concerns observations, not a demand that the two categorical variables already be independent. Testing whether variables are independent would be pointless if independence between them had to be known in advance.
Adequate expected counts matter for the usual chi-square approximation, and OpenStax presents a five-per-cell rule for its introductory tests, so sparse categories may need justified combining or a different method, not simply ignoring the low counts. A significant test may detect a minor practical difference in a large sample, so compare proportions and business importance as well as the p-value.
A small sample may miss a meaningful difference. Do not force a causal story onto an association: product type and delinquency may both reflect borrower risk or selection policies, and the chi-square statistic tests a pattern in category counts, while causal inference needs a separate design.
In practice
Real-world examples.
Example
A lender expects 50 defaults in each of two otherwise equally sized categories but observes 60 and 40. The statistic is 100/50 plus 100/50, or 4, before assessing degrees of freedom and a p-value.
Example
A card issuer builds a two-by-two table of product type and whether a complaint occurred. It compares each observed cell count with the count implied by independent classifications.
Example
An analyst obtains a small p-value from a huge dataset, then finds that the difference in complaint rates is only a fraction of a percentage point. Statistical and operational importance differ.
Formula
Calculation
Chi-square = sum over cells of (observed count - expected count)^2 / expected count. For independence, expected cell count = row total x column total / grand total. If a 2 x 2 table has row totals 80 and 120, column totals 50 and 150, and overall total 200, the first expected count is 80 x 50 / 200 = 20. Use all cells and the appropriate degrees of freedom to assess the full statistic.Case study
Seen in the real world.
Fictional example: A lender tests whether delinquency status is independent of two loan categories. Its table contains 200 distinct loans, and it calculates each expected count from the row and column totals. One cell's expected count is too small for the introductory approximation it planned to use.
The analyst does not publish a p-value from the unsuitable calculation. She checks whether a defensible grouping or another test can handle the sparse cells. In her report she distinguishes a possible association from proof that loan category itself caused delinquency, and she discloses how the sample was selected.
Watch out
Common mistakes.
- Entering percentages or continuous measurements into a categorical-count test as though they were independent category frequencies.
- Reading a small p-value as the probability a hypothesis is false or as proof of a meaningful effect.
- Assuming the test's observations must be independent of the two variables being compared, rather than checking independence among observations.
Questions
People also ask.
What does a large chi-square statistic mean?
The observed counts differ more from those expected under the specified null model. Degrees of freedom and assumptions determine the strength of evidence.
Does the test show causation?
No. A test of independence can support evidence of association, but cannot identify a causal mechanism by itself.
Can I use it with a very small expected count?
The usual approximation may be unsuitable. Consider a justified category grouping or an appropriate alternative test.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%