SEARCH · ENGLISH EDITION
Find an AI concept.
Search titles, summaries, article text, source-language titles, and known aliases.
EN
Benchmark Data Contamination
The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.
EN
HELM
Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.
EN
MMLU
Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.
EN
Safety Evaluation
Testing designed to measure risks, harmful behaviors, and the effectiveness of safeguards in an AI system.