22 terms
Safety & Ethics
Alignment, security, bias, and the risks people argue about most.
- AI safetynountrendingthe field concerned with making AI systems behave as intended and avoid causing harm, from everyday mistakes and misuse to risks…
- alignmentnountrendingthe work of making an AI system's goals and behavior match what its designers and users actually intend, including human values
- biasnounsystematic unfairness in an AI system's outputs, often learned from patterns in its training data, that favors or disadvantages…
- black boxnouna system whose inputs and outputs can be seen but whose internal reasoning cannot easily be understood, a common description of…
- data voidnouna topic with little reliable information online, so search engines and AI models fall back on whatever low-quality or misleading…
- deepfakenountrendinga realistic fake image, video, or audio recording of a real person, made with AI
- existential risknounthe possibility that advanced AI could cause human extinction or permanently and drastically limit humanity's future
- guardrailsplural nounrules, filters, and checks placed around an AI system to keep its outputs and actions within safe and acceptable limits
- interpretabilitynounresearch into understanding what is happening inside a model, such as which internal features represent which concepts, so its…
- jailbreaknouna prompt or trick designed to get an AI model to ignore its safety rules and produce content it is meant to refuse
- LLM groomingnoundeliberately flooding the internet with false or slanted content in the hope that AI models will absorb it during training or…
- model cardnouna document published alongside a model describing what it is for, how it was trained and tested, its limitations, and known risks
- neuralesenounthe internal numerical 'language' a model might use to think and communicate, which humans cannot read; often discussed as a risk…
- opaque recurrencenouna technique in which a model reasons by looping information through its internal layers several times instead of writing out its…
- p(doom)noun(informal) a person's estimated probability that AI will lead to catastrophic outcomes for humanity
- prompt injectionnountrendingan attack in which hidden instructions are slipped into content an AI reads, such as a web page, email, or document, causing it…
- recursive self-improvementnouna scenario in which an AI system improves its own design, and each improved version gets better at improving itself, potentially…
- red teamingnoundeliberately attacking or stress-testing an AI system, often with outside experts, to find harmful behaviors, security holes, and…
- responsible AInounthe practice of designing, building, and deploying AI with attention to fairness, transparency, privacy, safety, and…
- reward hackingnounwhen an AI finds a shortcut that earns a high score from its reward or test without actually doing the intended task
- sycophancynountrendingan AI model's tendency to tell users what they want to hear, agreeing with them, praising their ideas, or changing correct…
- watermarkingnounhiding an invisible signal in AI-generated text, images, audio, or video so it can later be identified as machine-made
Advertisement