18 terms
Training
How models are built and improved, from pre-training to fine-tuning and reinforcement learning.
- computenountrendingcomputing power, especially the chips and processing time used to train and run AI models
- constitutional AInouna training approach in which a model is guided by a written set of principles (a "constitution") and uses AI feedback, rather…
- data labelingnounthe work of tagging raw data with correct answers or categories, such as marking objects in photos or rating chatbot replies, so…
- distillationnountrendingtraining a smaller "student" model to imitate the outputs of a larger "teacher" model, producing a cheaper model that keeps much…
- emergent abilitiesplural nounskills that seem to appear suddenly once a model passes a certain size, even though it was not specifically trained for them
- epochnounone complete pass through the entire training dataset
- fine-tuningnountrendingfurther training of an existing model on a smaller, focused dataset so it gets better at a particular task, style, or domain
- instruction tuningnounfine-tuning a base model on examples of instructions paired with good responses, so it learns to follow requests instead of just…
- LoRAnounlow-rank adaptation: a cheap fine-tuning method that freezes the original model and trains only a small add-on, so a new style or…
- model collapsenouna gradual loss of quality and diversity that happens when models are trained on content produced by earlier models, so errors and…
- post-trainingnountrendingeverything done to a model after pre-training to make it useful and safe, such as instruction tuning, reinforcement learning, and…
- pre-trainingnounthe first and most expensive stage of building a model, in which it learns general patterns from an enormous amount of unlabeled…
- quantizationnounshrinking a model by storing its numbers with less precision (for example 4 bits instead of 16), which makes it smaller and…
- reinforcement learningnountrendinga way of training in which a model tries actions, receives rewards or penalties, and gradually learns which behavior earns the…
- reward modelnouna separate model trained to score how good a response is, used to guide reinforcement learning without a human rating every…
- RLHFnounreinforcement learning from human feedback: a training method in which people rate or compare a model's answers, and the model is…
- scaling lawsplural nountrendingobserved patterns showing that model performance improves in a predictable way as you increase model size, training data, and…
- synthetic datanountrendingtraining data generated by AI models or simulations rather than collected from the real world
Advertisement