LLM preference alignment algorithms

Large Language Model (LLM) preference alignment is a technical discipline focused on bridging the "alignment gap," which is the mismatch between static mathematical signals and the dynamic, multi-dimensional nature of human experiential meaning Verified Answer #2 Verified Answer #7. As of 2026, the field has matured from simple scalar reward modeling into "Reward Engineering," incorporating frameworks that account for human cognitive limitations and scalable oversight Verified Answer #4.

Core Alignment Paradigms

Reinforcement Learning from Human Feedback (RLHF)

RLHF is the foundational industry standard for encoding human values into models by relying on subjective human judgment rather than formal mathematical definitions Verified Answer #3.

Direct Preference Optimization (DPO)

DPO simplifies the alignment process by mapping the reward function directly to the policy, bypassing the need for a separate reward model or online reinforcement learning Verified Answer #1 Verified Answer #6.

Kahneman-Tversky Optimization (KTO)

KTO is based on principles of behavioral economics, specifically Prospect Theory, to align models using "human-aware loss functions" (HALOs) Verified Answer #7.

Advanced Alignment Frameworks

Cognitive-Aware and Game-Theoretic Models

Structural Decomposition

To avoid the pitfalls of lossy scalarization, some methodologies decompose complex human meaning into structured components Verified Answer #2.

Scalable Oversight

As models transition into autonomous agent ecosystems, recursive frameworks like Recursive Reward Modeling (RRM) use the model's own capabilities to assist in its evaluation, addressing the bottleneck of human-in-the-loop oversight Verified Answer #4.