ORBiS Tree

TOPIC · ENGLISH EDITION

Research

Research ideas, benchmarks, documentation, and evaluation.

EN

Benchmark Data Contamination

The presence of evaluation examples or closely related material in training data, potentially inflating benchmark results.

EN

Data Provenance

Information about where data came from, how it was collected, and how it changed before use.

EN

Datasheets for Datasets

A documentation framework for describing how datasets were created, composed, maintained, and intended to be used.

EN

Foundation Model

A broadly trained model that can be adapted to many downstream tasks and applications.

EN

HELM

Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.

EN

Hallucination in AI

A generated statement that is unsupported, fabricated, or inconsistent with the available evidence or context.

EN

MMLU

Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.

EN

Model Card

A structured document that describes a model's intended use, evaluation, limitations, and other important context.

EN

Red Teaming

Adversarial testing intended to discover failure modes, unsafe behavior, or exploitable weaknesses in an AI system.

EN

Safety Evaluation

Testing designed to measure risks, harmful behaviors, and the effectiveness of safeguards in an AI system.

EN

System Card

Documentation describing a deployed AI system, including capabilities, evaluations, mitigations, and limitations.