ORBiS TREE · ENGLISH DOCUMENT
Research →MMLU
Massive Multitask Language Understanding, a benchmark covering many academic and professional subject areas.
한국어
English
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
MMLU is a benchmark designed to test knowledge and problem solving across a broad collection of subjects. It contains multiple-choice questions spanning areas such as mathematics, law, history, computer science, and medicine.
How it works
Models are evaluated by selecting answers under a defined prompting and scoring procedure. Reported scores can vary with prompt format, few-shot examples, model version, contamination, and evaluation implementation.
Why it matters
MMLU became a widely cited general capability benchmark, but it should not be treated as a complete measure of intelligence or product quality. Broader evaluation suites and task-specific tests remain necessary.
Related concepts
SOURCES
Sources
-
Measuring Massive Multitask Language UnderstandingHendrycks et al. / ICLR / arXivOpen source ↗
KNOWLEDGE LINKS
Continue from here
BACKLINKS