Back to Glossary

Entry · Business

Fraud Risk Score

A fraud risk score is a model or rules engine's estimate or ranking of how suspicious a transaction, account or action is. Its scale and direction are vendor-specific; a score alone does not prove fraud or dictate one universal response.

From the Money Master HQ dictionary, founded by Shihan Sheriff (FCMA, VP of Finance at Nomod, CFO at Esanjo Ventures). How these definitions are written.

What it means

Payment and account systems process too many events for every one to receive manual review, so a score helps sort cases for approval, extra checks, review or blocking. It is one input in a risk decision.

Stripe documents risk evaluation for card payments, and Adyen documents machine-learning risk rules, but their scoring and action designs differ, so do not assume that every system uses 0 to 100 or that a high number always means high risk. A fictional online shop receives a risk score on each payment, and its operations team reads the provider's scale before setting a threshold, because copying another provider's settings would be unsafe.

Models can use transaction amount, device signals, location, history and behaviour where available and lawful, though the exact features and weights may be proprietary and some signals can be missing or unreliable. A fictional cardholder who buys from a new device while travelling may raise the score, yet the purchase could still be legitimate.

Scores may be probabilistic estimates, relative rankings or labels mapped to numbers, so calibration matters: a value of 80 is not automatically an 80% chance of fraud, and the provider documentation should be read. A fictional analyst labels 80 as "high" in one tool while in another tool 80 indicates low risk, so the team records the meaning in its operating procedure.

A threshold determines what happens next, since businesses might allow low-risk activity, request authentication for a middle band and review or decline high-risk events, and these are design choices, not a universal formula. A fictional merchant tests a stricter blocking threshold, and fraud losses fall but legitimate customers are declined more often, so it weighs both outcomes before keeping the change.

False positives are genuine customers incorrectly flagged, while false negatives are harmful activity missed, and a useful system seeks a workable balance rather than the lowest possible fraud count. A fictional subscription service that manually reviews every flagged new signup delays genuine users, so it changes its workflow after measuring approval and complaint rates.

Track outcomes using reliable labels such as confirmed chargebacks or investigated account abuse, with suitable delay, because a disputed charge is not automatically proven fraud and labelling errors can distort model evaluation. A fictional team scores payments in January, then reviews later confirmed outcomes, avoiding judging the model solely from same-day disputes.

Risk patterns also shift, as a holiday surge, new market or attacker method can make old thresholds less useful, so monitor data drift and review decisions periodically, as a fictional retailer does when a launch in a country with unfamiliar payment patterns flags more orders than usual and the team checks legitimate approval impact before changing thresholds. Explain decisions to customers as clearly as security permits, since a generic "payment failed" response may send people into repeated attempts and support staff need a safe escalation path; a fictional buyer wrongly blocked is handled through an approved review route, not by asking for sensitive card details by email, and the risk score is not disclosed as proof of wrongdoing.

Privacy and fairness matter, so use permitted data, control access and review whether rules disproportionately burden groups or regions, with compliance requirements depending on the product and jurisdiction, and a fictional team seeing high reviews from a particular region examines data quality and outcomes instead of automatically blocking that geography. A score should complement authentication, velocity checks, staff review and dispute analysis because attackers adapt to single thresholds, and the team should document provider version, score meaning, thresholds and measured trade-offs, since the number is a decision aid, not a verdict on a person.

In practice

Real-world examples.

1

Example

A merchant routes medium-risk payments to authentication.

2

Example

A score spike during travel is reviewed without assuming fraud.

3

Example

A business measures legitimate declines after tightening thresholds.

Formula

Calculation

No universal score formula. Evaluate a chosen threshold with confirmed fraud caught, legitimate activity declined and review workload on the same defined sample. Worked comparison on a fictional sample of 10,000 payments, of which 200 are later confirmed as fraud and 9,800 are legitimate. Threshold A blocks 258 payments: 160 confirmed fraud and 98 legitimate. It catches 160 / 200 = 80% of fraud and declines 98 / 9,800 = 1% of legitimate payments. Stricter threshold B blocks 425 payments: 180 confirmed fraud and 245 legitimate. It catches 180 / 200 = 90% of fraud but declines 245 / 9,800 = 2.5% of legitimate payments. Moving from A to B catches 20 more fraudulent payments but declines 147 more legitimate ones (245 - 98). At an average order of $80, that is about 20 x $80 = $1,600 of extra fraud prevented against 147 x $80 = $11,760 of legitimate sales declined, before chargeback fees, review costs and customer goodwill. The stricter threshold is not automatically better.

Case study

Seen in the real world.

In this fictional case, Lumen Shop blocks payments above a provider-specific threshold. Chargebacks fall, but completed orders also drop. The team audits false positives, tests authentication for a middle band and documents the provider's score scale before adjusting the rule. After three months, the fictional team reviews the same measures again on a fresh sample. It finds that the middle-band authentication step recovers some orders that had previously been declined, but a few customers abandon at the extra step, so it keeps the change under review rather than declaring the problem solved.

Watch out

Common mistakes.

  • Assuming every score uses the same scale and direction.
  • Treating a high score as proof of fraud.
  • Optimising fraud loss without counting legitimate declines.

Questions

People also ask.

Does a score of 80 mean an 80 percent fraud chance?

Not necessarily. Check the provider's score definition and calibration.

Should every high-scoring payment be blocked?

Not automatically. Review the business's risk, false positives and available checks.

Why can a legitimate customer be flagged?

New devices, unusual behaviour or poor data may resemble risky patterns.

Was this explanation helpful?

From the founder's library

Accounting Fundamentals: A Non-Finance Manager's Guide to Finance and Accounting, by Shihan Sheriff

Take it further with the book.

Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.

US$2.24US$2.99

25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.

View the book and save 25%
Last updated · October 8, 2026
Browse all terms →

Disclaimer

The information provided in this finance dictionary is for educational and informational purposes only. It should not be construed as financial, investment, legal, or tax advice. Always consult with a qualified professional before making any financial decisions. Money Master HQ makes no representations or warranties about the accuracy, completeness, or suitability of this information. Use of this content is at your own risk.