ORBiS Tree

ORBiS TREE · ENGLISH DOCUMENT

Research →

Safety Evaluation

Testing designed to measure risks, harmful behaviors, and the effectiveness of safeguards in an AI system.

한국어 English
English updated 8월 29, 2026 Source article updated 8월 29, 2026 1 sources
This English page is a curated translation layer linked to the Korean source article. Community changes are currently made on the Korean source, where the full revision history and anonymous edit trail are preserved.
Safety evaluation examines whether an AI system exhibits behaviors that could cause harm under realistic or adversarial conditions. The relevant risks depend on the model, tools, deployment environment, and user population.

How it works

Evaluations may measure policy violations, deception, bias, privacy leakage, cyber or biological misuse potential, robustness, tool-use failures, or other domain-specific risks. Good evaluation defines clear scenarios, metrics, thresholds, and model versions.

Why it matters

Safety evaluation is an ongoing process rather than a one-time certification. New capabilities, prompts, tools, and deployment contexts can create new failure modes after an initial release.

Related concepts

SOURCES

Sources

  1. Open source ↗