ORBiS TREE · ENGLISH DOCUMENT
Research →Data Provenance
Information about where data came from, how it was collected, and how it changed before use.
한국어
English
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
Data provenance records the origin and processing history of data. In AI systems it can include source datasets, collection methods, licenses, filtering, transformations, deduplication, and version history.
How it works
Reliable provenance requires traceable identifiers and documentation across the data pipeline. It becomes harder when training corpora combine many sources or when intermediate datasets are repeatedly transformed.
Why it matters
Provenance supports reproducibility, copyright and license review, quality analysis, bias investigation, and contamination detection. It is therefore closely connected to governance as well as model performance.
Related concepts
SOURCES
Sources
-
The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AIData Provenance Initiative / arXivOpen source ↗
KNOWLEDGE LINKS
Continue from here
BACKLINKS