Making models behave as intended.
THE DEFINITION
Additional training so a model follows instructions, is helpful, and declines harmful requests.
THE PROBLEM IT ADDRESSES
A model trained purely to predict text produces whatever the training data suggests.
HOW IT IS DONE
Training on examples of good responses Human feedback ranking outputs Explicit rules and guidelines
WHAT IT PRODUCES
Models that answer questions rather than continuing your text Models that decline certain requests A consistent tone
WHAT IT DOES NOT SOLVE
Factual accuracy. An aligned model is helpful and still hallucinates.
A SIDE EFFECT
Models trained toward helpfulness tend to agree rather than challenge.
Ask explicitly for criticism.
WHY MODELS SOMETIMES REFUSE REASONABLE REQUESTS
Alignment involves judgement applied broadly. Some legitimate requests resemble problematic ones.
RELATED TERMS
Hallucination, refusal, sycophancy.