Skip to content
Machine Learning Glossary

RLHF

Reinforcement Learning from Human Feedback, a method that aligns models with human preferences using ranked examples.

Last updated

RLHF trains a model to produce answers people prefer. Humans rank candidate responses, a reward model learns those preferences, and reinforcement learning optimises the assistant against it. This is a key step behind the helpful, conversational tone of modern chat models, and it is complemented today by related techniques such as direct preference optimisation.

Related terms

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page