ORBiS TREE · ENGLISH DOCUMENT
Research →Benchmark Data Contamination
The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.
한국어
English
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
Benchmark data contamination occurs when test items, answers, or highly similar material appear in a model's training data. This can make a benchmark score look better without representing genuine generalization.
How it works
Contamination can arise from public benchmark datasets, scraped web pages, derivative datasets, or repeated discussions of test items. Detecting it may require exact matching, similarity analysis, provenance records, and controlled evaluation sets.
Why it matters
As models are trained on increasingly broad web corpora, contamination has become a central concern in interpreting benchmark claims. Transparent evaluation should disclose known risks and use fresh or protected tests where possible.
Related concepts
SOURCES
Sources
KNOWLEDGE LINKS
Continue from here
BACKLINKS