LLM reasoning / chain-of-thought
SPACE-3: Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation
Superseded — cited as a baseline and beaten by newer methods
0 papers critique it · 1 beat it on benchmarks
Head-to-head results where a newer method reports beating SPACE-3. Values are copied from the source paper's tables — verify against the cited paper.
GEM (BERT_base + GNN) / T5_small + Claude 3.7 S. beats SPACE-3
65.19 vs 57.50
JGA · [Multi-domain DST on MultiWOZ 2.2]