UPDATED 2026-08-04
RLHF (Reinforcement Learning from Human Feedback)
What RLHF means and why AI labs hire human evaluators.
RLHF uses human preference data to fine-tune AI models. Evaluators compare model outputs and select which is better, safer, or more helpful — work that is often paid as remote contracts.