LLM reasoning / chain-of-thought

PPO (Proximal Policy Optimization)

Superseded baseline#532 of 772 most-superseded

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites PPO (Proximal Policy Optimization) as a baseline.

Foundational approaches like Reinforcement Learning from Human Feedback (RLHF), typically implemented via Proximal Policy Optimization (PPO), are powerful but suffer from high computational overhead and hyperparameter sensitivity.
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.