Mixture-of-experts routing

DPO

Superseded baseline#63 of 1,370 most-superseded

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 1 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites DPO as a baseline.

requires explicit human or teacher model judgments to label one response as preferred over another, making them costly and hard to scale
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
standard DPO methods are still inherently limited by their reliance on a single, monolithic policy.
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.