Machine Learning
Glossary
RLHF
Reinforcement Learning from Human Feedback, a method that aligns models with human preferences using ranked examples.
Last updated
RLHF trains a model to produce answers people prefer. Humans rank candidate responses, a reward model learns those preferences, and reinforcement learning optimises the assistant against it. This is a key step behind the helpful, conversational tone of modern chat models, and it is complemented today by related techniques such as direct preference optimisation.
Related terms
Keep exploring
Browse the full AI glossary or compare AI tools that use this technology.