ORBiS Tree

ORBiS TREE · ENGLISH DOCUMENT

Research →

HELM

Holistic Evaluation of Language Models, a framework for evaluating models across scenarios and metrics.

한국어 English
English updated 8월 29, 2026 Source article updated 8월 29, 2026 1 sources
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
HELM is an evaluation framework developed to make language-model comparisons more systematic across tasks, metrics, and deployment concerns. It emphasizes that a single accuracy number does not capture all relevant behavior.

How it works

The framework evaluates models across scenarios and can consider dimensions such as accuracy, calibration, robustness, fairness, efficiency, and other properties depending on the release and configuration.

Why it matters

HELM is useful as an example of multi-dimensional evaluation. It reinforces the need to document prompts, model versions, datasets, and metrics when comparing general-purpose models.

Related concepts

SOURCES

Sources

  1. Open source ↗