What it means
The label covers two quite different situations. Some models are secret by choice, such as a vendor's proprietary scoring system, while others are open but genuinely uninterpretable, such as a neural network with millions of parameters whose behaviour no one can summarise in a sentence.
Businesses adopt them because they often work better. A model that combines hundreds of weak signals will usually out-predict a simple rule set, and in areas like fraud detection or advertising bidding that accuracy translates directly into money saved or earned.
The problem is accountability. When a customer is refused credit, a regulator asks why a price was set, or a trading strategy loses money for six months, "the model said so" is not an acceptable answer, and firms need to be able to explain and defend the decision.
Model risk management has therefore become a discipline of its own. Standard practice includes independent validation before deployment, testing on data the model has never seen, running a simpler challenger model alongside it for comparison, monitoring for drift as conditions change, and documenting the limits of where the model may be applied.
Explainability tools help without solving the problem entirely. Techniques that show which inputs mattered most for a particular decision give a usable narrative for customers and regulators, though they approximate the model's reasoning rather than reveal it.
The judgement call for management is where opacity is acceptable. A black box that recommends which advert to show carries little consequence if it is wrong, while one that decides mortgage approvals or capital requirements needs far more scrutiny and often a simpler, explainable alternative.
In practice
Real-world examples.
Example
A consumer lender buys a third-party scoring model that improves approval accuracy but cannot explain individual decisions. To meet its obligation to give reasons for a refusal, it runs a simple, transparent scorecard alongside it and uses that for customer communication. Where the two disagree sharply, a credit officer reviews the case by hand before any decision is issued.
Example
A quantitative fund runs a strategy generated by machine learning across hundreds of signals. Investors accept the opacity but insist on strict position limits and a hard stop-loss, so that a model failure cannot become an existential loss. The fund also publishes monthly attribution showing which broad factors drove returns, which reassures allocators without revealing the strategy itself.
Example
An insurer prices motor cover using a complex model that a regulator asks it to justify. The insurer produces a factor-level analysis showing which variables drive premiums and demonstrates that protected characteristics are neither used nor closely proxied by other inputs. That evidence pack now forms part of every annual pricing review rather than being assembled only when a regulator asks.
Case study
Seen in the real world.
Merrifield Credit is a fictional lender used here as an illustrative example. It replaced its long-standing scorecard with a purchased machine-learning model that promised a meaningful lift in approval accuracy on the same volume of applications, and approvals rose immediately without any obvious deterioration in early arrears.
Eighteen months later, defaults on the $400 million book had risen from 3.1% to 5.4%, an extra 2.3 percentage points, or roughly $9.2 million of additional losses. Because nobody at Merrifield could explain how the model weighted its inputs, it took a further four months to establish that a data feed had changed definition and that the model had been quietly relying on it.
The illustrative lesson is not that the model was bad, but that Merrifield had no way to interrogate it. After the review the lender kept the model but added a transparent challenger scorecard, monthly monitoring of input distributions, and a documented rule that any divergence beyond a set threshold would trigger a manual review.
Watch out
Common mistakes.
- Assuming that high accuracy on historic data means a model will hold up in future conditions it has never encountered.
- Buying a vendor model without contractual rights to validation evidence, monitoring data or an explanation of its inputs.
- Confusing a model that is merely complicated with one that is genuinely unexplainable, and giving up on scrutiny too early.
Questions
People also ask.
Are black box models banned in regulated decisions?
Generally no, but rules on explaining automated decisions and on fair treatment mean a firm must be able to justify outcomes whatever technique produced them.
How do you test a model you cannot see inside?
By testing behaviour rather than internals: hold-out data, stress scenarios, sensitivity to each input, and comparison with a simpler benchmark.
What is model drift?
It is the gradual decay of a model's accuracy as the world changes away from the data it was built on, which is why continuous monitoring matters more than a single sign-off.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%