llm-preference-alignment-algorithms

Large language model (LLM) preference alignment aims to bridge the "alignment gap," which is the mismatch between static mathematical signals and the multi-dimensional nature of human experiential meaning Verified Answer #1. Modern alignment research has shifted from monolithic scalar reward modeling toward frameworks that decompose, contextualize, and direct preference learning Verified Answer #1.

Core Alignment Paradigms

Reinforcement Learning from Human Feedback (RLHF)

Direct Preference Optimization (DPO)

Kahneman-Tversky Optimization (KTO)

Reward Engineering and Structural Decomposition

Reward Engineering is an emerging discipline that moves beyond standard preference modeling to account for human cognitive limitations and scalable oversight Verified Answer #4.

Cognitive-Aware Modeling

Rule-Based and Rubric Frameworks

Scalable Oversight

Scalable oversight is a recursive framework designed to overcome the bottleneck of human-in-the-loop evaluation as models transition into autonomous agent ecosystems Verified Answer #4. This architecture utilizes the model's own capabilities to assist in its own evaluation, with Recursive Reward Modeling (RRM) serving as a primary mechanism Verified Answer #4.