ORBiS Tree

ORBiS TREE · ENGLISH DOCUMENT

Research →

Benchmark Data Contamination

The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.

한국어 English
English updated 8월 29, 2026 Source article updated 8월 29, 2026 1 sources
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
Benchmark data contamination occurs when test items, answers, or highly similar material appear in a model's training data. This can make a benchmark score look better without representing genuine generalization.

How it works

Contamination can arise from public benchmark datasets, scraped web pages, derivative datasets, or repeated discussions of test items. Detecting it may require exact matching, similarity analysis, provenance records, and controlled evaluation sets.

Why it matters

As models are trained on increasingly broad web corpora, contamination has become a central concern in interpreting benchmark claims. Transparent evaluation should disclose known risks and use fresh or protected tests where possible.

Related concepts

SOURCES

Sources

  1. Open source ↗