LLM reasoning / chain-of-thought
Block Universal Transformer
Superseded baseline#36 of 772 most-superseded
Superseded — cited as a baseline and beaten by newer methods
1 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Block Universal Transformer as a baseline.
Flattening the architecture by reusing a single Transformer block yields modest gains over the standard Transformer, but overall performance remains unsatisfactory.
Beaten on benchmarks
Head-to-head results where a newer method reports beating Block Universal Transformer. Values are copied from the source paper's tables — verify against the cited paper.
SR² beats Block Universal Transformer
93.7 vs 30.4
Maze-Hard · [Maze-Hard]
Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.