Back to Glossary

Entry · Business

Data Lake

A data lake is a storage environment designed to hold large volumes of data in varied formats, often in its original form until a particular use is defined. It can support analytics, machine learning and data sharing across sources. Its flexibility only helps when information is catalogued, protected and usable; collecting files without governance can create a hard-to-search data swamp.

From the Money Master HQ dictionary, founded by Shihan Sheriff (FCMA, VP of Finance at Nomod, CFO at Esanjo Ventures). How these definitions are written.

What it means

A retailer has sales tables, website events, product images and service transcripts, and a traditional report database may require each source to be reshaped before storage. A data lake can retain these varied inputs and prepare them later for chosen analyses.

IBM describes data lakes as repositories for large volumes of raw structured, semi-structured and unstructured data, often on scalable object storage, and AWS describes the same broad ability to store original data and use it for different analyses. Neither implies that every lake is automatically clean or safe, so identify what data belongs in the lake and why, because "store everything" can create cost and privacy risk without useful outcomes.

Plan ingestion, since data may arrive from applications, databases, devices or partners on different schedules, and preserve source information and timestamps so analysts know where a file came from and when it was captured. A raw zone can keep an original copy, while curated zones may clean, validate and organise data for reliable use.

Specify access rules for each zone, because a copy of payroll data does not become public merely because it sits next to product images. Create a catalogue of names, owners, formats and meanings, since without metadata a team may spend more time hunting for usable data than analysing it.

Decide retention too, as keeping duplicates and obsolete exports forever increases storage and security burden. A lake can store data before a fixed reporting schema is chosen, and analysts may apply structure when reading, but that is not permission to skip validation.

A data warehouse generally focuses on curated data for known reporting needs, and the two can complement each other. A lakehouse combines lake storage with features for more managed analytical tables, but product labels can overlap, so inspect actual capabilities.

Track data quality, because missing identifiers, duplicate events and inconsistent currencies can still distort a dashboard, and version transformation logic so that a report using changed rules is not silently compared with an older one. Control compute costs as well as storage costs, since large scans can become expensive even if storing files is cheap, and monitor ingestion failures because a quiet pipeline can make a daily report look stable while yesterday's data is missing.

Review lifecycle rules as source systems change, as a feed can change format without warning and break downstream work. Protect personal and regulated records according to current law, contract and organisational policy; encryption, access control and logging may be needed, but a single technical control is not sufficient by itself.

Define an owner for the lake platform and owners for the data domains within it, separate development from production where appropriate, check how data leaves the lake, and plan backup and recovery according to whether data can be reproduced from sources. For a new use case, start with a bounded dataset and a clear decision, and measure value through useful reports and reduced duplication, not terabytes stored, because a data lake is an architecture option, not a guarantee that a business has become data-driven.

In practice

Real-world examples.

1

Example

A company stores click events, product pictures and sales tables in one governed environment. Each source has a named owner and a catalogue entry describing its format. Raw files are kept as received while curated tables are built for reporting.

2

Example

An analyst uses a curated sales table while the original event files remain in a raw zone. When a number looks odd, the analyst can trace it back to the source file. The raw zone is not opened to general users.

3

Example

A daily feed fails, and a freshness alert prevents yesterday's incomplete dashboard from being used. The data engineer reruns the load, and the dashboard is flagged as delayed until the missing data arrives. Managers do not make decisions on partial figures.

Formula

Calculation

No universal formula. A useful freshness check is datasets delivered on schedule / datasets expected on schedule x 100, supplemented by quality checks. Worked example. A lake expects 40 daily feeds and 38 arrive on schedule, so freshness is 38 / 40 x 100 = 95%. The two late feeds, or 5%, are investigated with their owners, and the dashboards that depend on them are marked delayed until the data arrives.

Case study

Seen in the real world.

This entirely fictional case follows Willow Markets. Its first lake collected sales and web events, but teams could not tell which files were current. The company added owners, a catalogue and freshness alerts before building a cross-channel report. It removed unused exports under a retention policy.

The case is invented. Willow then measured freshness: of 40 expected daily feeds, 38 arrived on schedule, or 95%, and the two late feeds were investigated with their owners. Retention rules removed duplicate exports that no report used. The figures are illustrative.

Watch out

Common mistakes.

  • Treating cheap storage as permission to keep all data forever.
  • Assuming raw data is ready for financial reporting without validation.
  • Forgetting access controls on copied sensitive records.

Questions

People also ask.

Is a lake the same as a warehouse?

No. Lakes commonly hold varied raw data; warehouses emphasise curated data for known reporting uses.

Does schema-on-read eliminate data quality work?

No. It delays some structure choices, but validation remains necessary.

Can a small business use one?

Yes if a concrete use justifies the operational cost and governance work.

Was this explanation helpful?

From the founder's library

Accounting Fundamentals: A Non-Finance Manager's Guide to Finance and Accounting, by Shihan Sheriff

Take it further with the book.

Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.

US$2.24US$2.99

25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.

View the book and save 25%
Last updated · October 8, 2026
Browse all terms →

Disclaimer

The information provided in this finance dictionary is for educational and informational purposes only. It should not be construed as financial, investment, legal, or tax advice. Always consult with a qualified professional before making any financial decisions. Money Master HQ makes no representations or warranties about the accuracy, completeness, or suitability of this information. Use of this content is at your own risk.