ORBiS Tree

ORBiS TREE · ENGLISH DOCUMENT

Research →

MMLU

Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.

한국어 English
English updated 8월 29, 2026 Source article updated 8월 29, 2026 1 sources
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
MMLU is a benchmark designed to test knowledge and problem solving across a broad collection of subjects. It contains multiple-choice questions spanning areas such as mathematics, law, history, computer science, and medicine.

How it works

Models are evaluated by selecting answers under a defined prompting and scoring procedure. Reported scores can vary with prompt format, few-shot examples, model version, contamination, and evaluation implementation.

Why it matters

MMLU became a widely cited general capability benchmark, but it should not be treated as a complete measure of intelligence or product quality. Broader evaluation suites and task-specific tests remain necessary.

Related concepts

SOURCES

Sources

  1. Open source ↗