SEARCH · ENGLISH EDITION
Find an AI concept.
Search titles, summaries, article text, source-language titles, and known aliases.
Benchmark Data Contamination
The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.
Claude
A family of general-purpose AI models and conversational products developed by Anthropic.
Data Provenance
Information about where data came from, how it was collected, and how it changed before use.
Datasheets for Datasets
A documentation framework for describing how datasets were created, composed, maintained, and intended to be used.
HELM
Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.
Hallucination in AI
A generated statement that is unsupported, fabricated, or inconsistent with the available evidence or context.
Latency and Throughput
Two core serving metrics describing response delay and the amount of work a system completes over time.
MMLU
Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.
Red Teaming
Adversarial testing intended to discover failure modes, unsafe behavior, or exploitable weaknesses in an AI system.