What it means
The starting point is a proposed explanation for the data, which might specify how frequently different outcomes should occur or which distribution describes a measurement. Observed results are then compared with what that model would predict.
A test needs a null hypothesis: in a chi-square goodness-of-fit test, the null typically says that the data follow the specified distribution or category proportions, and the alternative says they do not, subject to the way the test has been constructed. Expected values should come from the stated model rather than be chosen after seeing the results to make them fit.
If model parameters are estimated from the same data, the test's degrees of freedom may need adjustment, and NIST's guidance explains the role of estimated parameters in the chi-square calculation. The chi-square method adds squared differences between observed and expected counts, scaled by expected counts, so large discrepancies increase the statistic.
A suitable reference distribution and degrees of freedom are then used to interpret how unusual the result is under the null. The expected counts and other assumptions matter, since very small expected counts, unsuitable grouping or dependent observations can make a standard calculation unreliable.
A spreadsheet producing a number is not evidence that the statistical method is appropriate for the data. A p-value is not the probability that the model is true, because it describes the extremeness of results under the null assumptions.
A large p-value may reflect limited evidence against the model, including limited sample information, rather than proof of an accurate explanation. Statistical significance is also different from operational importance, and a smaller dataset may fail to detect a meaningful problem, so managers should consider effect size and decision consequences as well as the test result.
Goodness-of-fit is also different from testing independence between two categorical variables. Both can use chi-square calculations, but they ask different questions and have different expected-count structures, so use the method that matches the business question rather than choosing a familiar formula by name.
A model that fits existing data may still forecast poorly, since changes in customers, processes or economic conditions can alter the distribution. Validate important models on relevant additional data and monitor performance instead of relying forever on one successful fit assessment.
For managers, the useful output states the model, sample, assumptions, test result and practical discrepancy, and explains what action would follow if the model is unsuitable. This turns the test into evidence for a decision rather than a decorative statistical score in a report.
In practice
Real-world examples.
Example
A credit team expects applications to arrive in three categories with defined proportions. It compares observed counts with the expected counts to assess whether that allocation model fits the current sample. The team records the sample size and the proportions it assumed.
Example
An operations analyst obtains a small p-value for a minor difference in a very large dataset. The team assesses whether the difference changes staffing needs before redesigning the process. A statistically detectable gap of one percentage point may not justify any change.
Example
A manager interprets a non-significant result as proof that a forecast model is correct. The analyst explains that the test did not find enough evidence against the model under the stated assumptions. She recommends testing the model again on the next quarter's data.
Formula
Calculation
Chi-square statistic = sum of (observed count - expected count)^2 / expected count. For observed counts of 50, 30 and 20 against expected counts of 40, 35 and 25, the contributions are (50 - 40)^2 / 40 = 2.5, (30 - 35)^2 / 35 = about 0.714 and (20 - 25)^2 / 25 = 1, totalling about 4.214.
With three categories and no estimated parameters, the degrees of freedom are 3 - 1 = 2. A commonly published 5% critical value for 2 degrees of freedom is about 5.99, so a statistic of 4.214 would not be flagged as unusual at that level, although a sample of 100 is small and the result says little about practical importance.
Interpretation requires appropriate degrees of freedom and assumptions. The statistic alone is not a p-value, and the calculation should not be treated as a complete test without checking expected counts and whether parameters were estimated.Case study
Seen in the real world.
Fictional case study: Harbor Payments used a model predicting the mix of transaction types. A monthly report marked the model as reliable because a goodness-of-fit test had not rejected its expected proportions. The reviewer added the sample size, expected counts and test assumptions. The team also checked whether the operational differences were large enough to affect staffing and compared the pattern with later data.
Harbor retained the model for planning but removed the claim that the test proved it correct. It established a review process for changing transaction patterns, using the statistical result alongside practical evidence rather than as a permanent certificate of accuracy. Each quarter, the analyst reruns the comparison on fresh data and reports any category whose actual share has drifted from the modelled share.
Watch out
Common mistakes.
- Assuming goodness-of-fit applies only to normal distributions. Tests can assess other specified distributions or category proportions.
- Treating a large p-value as proof that the model is true. Failure to reject is not confirmation of perfect fit or future accuracy.
- Confusing goodness-of-fit with a test of independence. The business question and expected-count calculation differ.
Questions
People also ask.
Does a good fit guarantee a good forecast?
No. A model can fit past observations but fail when conditions change or when tested on different data.
Is the chi-square statistic itself a p-value?
No. Its interpretation depends on an appropriate reference distribution, degrees of freedom and valid assumptions.
What should a manager ask before using the result?
Ask which model was tested, what data were used, whether assumptions hold and whether the discrepancy matters for the decision.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%