What is data provenance?
Data provenance is the record of where a piece of data comes from. It names the source of a value, the person or program that entered it, and the time of the entry. With this record a company can check a figure before it relies on it.
What a provenance record holds
The W3C defines provenance as information about the entities, activities and people involved in producing a piece of data (W3C PROV overview). Its data model has three core types. An entity is the data itself, such as a record or a file. An activity is what happened to it, such as an import or an edit. An agent is the person, organization or program responsible for it.
For one value in a customer table, this comes down to four questions. Which source stated the value? Who entered it? When was it entered? From when was it true?
Most business databases cannot answer them. They keep one value per field and overwrite it when it changes. The old value is gone, and its source is gone with it.
Data provenance vs. data lineage
The two terms are often used for the same thing. Where they are kept apart, the difference is this:
| Data provenance | Data lineage | |
|---|---|---|
| Question | Where does this value come from? | Which path did the data take through the systems? |
| Looks at | One fact or one record | A table, a report or a pipeline |
| Holds | Source, author and dates of a value | Source systems, transformations and target tables |
| Typical use | An editor checks a value before it is published | A data engineer traces an error in a report |
Lineage shows that a column of a report was loaded from a table in the CRM. Provenance shows that the address in one CRM record came from an entry in the company register.
Why is data provenance important for AI?
An AI assistant that answers from the records of a company repeats what they say. It gives no hint when an address is outdated. A value with a source and a date can be judged by the person who reads the answer. Without them the reader has to trust the answer or look the value up again.
How a database keeps provenance
One way is to store a fact as a claim. The database then knows where a fact comes from and when it was true. A claim holds the value, the author, the source it cites and a trust rank. A correction adds a new claim and hides the old one. The database can still show what it said before.
A claim carries two dates. The first says when the fact was true. The second says when the database learned it. The technical term is bitemporal modeling. With both dates the store answers two questions: what was true on a given day, and what did we believe on that day.
Provenance needs a fixed object to point to. When two records describe the same company, their sources belong together. Master data management sets the rules for that, and object identity decides which records mean the same thing in the world.
My work on data provenance
I am an IT expert for data platforms. I have built databases and the tools around them for more than twenty years. I programmed the data stores of a global real estate database myself.
Today I work on a client’s information platform. The claims, the two dates and the trust rank are in production use there. I use W3C PROV-O as a reference and do not call the store compliant with it. A merge of two duplicates keeps the retired record as an alias, so a wrong merge can be taken back. Claude and Codex speed up this work under rules that a machine checks.
In data quality consulting I start with a measurement. One of its counts is how many records name no source and no date. I take on this work as a freelancer or as a permanent employee.
Searches this page answers
- what is data provenance
- data provenance and lineage tracking
- why is data provenance important
- data lineage
- data lineage vs audit trail
- data traceability
- provenance metadata
- bitemporal modeling
- what is bitemporal data
- valid time vs transaction time
- source reliability
- w3c prov
- w3c prov data model