SEARCH · ENGLISH EDITION
Find an AI concept.
Search titles, summaries, article text, source-language titles, and known aliases.
EN
RLHF
Reinforcement Learning from Human Feedback, a method for aligning model behavior with human preferences.
EN
Red Teaming
Adversarial testing intended to discover failure modes, unsafe behavior, or exploitable weaknesses in an AI system.
EN
Safety Evaluation
Testing designed to measure risks, harmful behaviors, and the effectiveness of safeguards in an AI system.
EN
System Card
Documentation describing a deployed AI system, including capabilities, evaluations, mitigations, and limitations.