LLM reasoning / chain-of-thought

DPO (Direct Preference Optimization)

Superseded baseline#262 of 772 most-superseded

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites DPO (Direct Preference Optimization) as a baseline.

While alternatives like Direct Preference Optimization (DPO) simplify the process by using static preference data, their efficacy is limited in tasks requiring dynamic interaction.
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.