SEARCH · ENGLISH EDITION
Find an AI concept.
Search titles, summaries, article text, source-language titles, and known aliases.
Benchmark Data Contamination
The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.
Data Provenance
Information about where data came from, how it was collected, and how it changed before use.
Datasheets for Datasets
A documentation framework for describing how datasets were created, composed, maintained, and intended to be used.
Fine-tuning
Additional training that adapts a pre-trained model to a narrower task, domain, or behavior.
HELM
Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.
Large Language Model (LLM)
A general-purpose language model trained on large datasets and usually built at substantial scale.
Model Card
A structured document that describes a model's intended use, evaluation, limitations, and other important context.
RLHF
Reinforcement Learning from Human Feedback, a method for aligning model behavior with human preferences.