Retrieval-augmented generation

RAFT

RAFT: A Real-World Few-Shot Text Classification Benchmark

Superseded baseline#31 of 1,179 most-superseded · first seen Sep 28, 2021

Superseded — cited as a baseline and beaten by newer methods

3 papers critique it · 4 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites RAFT as a baseline.

However, it suffers from conditional memorization bias and canonical answer overfitting.
Systematic Knowledge Injection into Large Language Models via Diverse Augmentation for Domain-Specific RAG
However, RAFT-trained models exhibit a critical limitation: they are conditioned to answer queries even when provided with entirely noisy contexts.
Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
RAFT solely focuses on identifying helpful information from retrieved documents. It learns to mimic the structured output format of teacher models that extract and directly quote sentences, rather than fostering domain thinking—unleashing reasoning capabilities involving higher-order cognitive processes.
RARE: Retrieval-Augmented Reasoning Modeling

Beaten on benchmarks

Head-to-head results where a newer method reports beating RAFT. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.