ORBiS Tree

SEARCH · ENGLISH EDITION

Find an AI concept.

Search titles, summaries, article text, source-language titles, and known aliases.

Results for “Benchmark”
EN

Benchmark Data Contamination

The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.

EN

Claude

A family of general-purpose AI models and conversational products developed by Anthropic.

EN

Data Provenance

Information about where data came from, how it was collected, and how it changed before use.

EN

Datasheets for Datasets

A documentation framework for describing how datasets were created, composed, maintained, and intended to be used.

EN

HELM

Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.

EN

Hallucination in AI

A generated statement that is unsupported, fabricated, or inconsistent with the available evidence or context.

EN

Latency and Throughput

Two core serving metrics describing response delay and the amount of work a system completes over time.

EN

MMLU

Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.

EN

Red Teaming

Adversarial testing intended to discover failure modes, unsafe behavior, or exploitable weaknesses in an AI system.