CLMar 26, 2024

Naive Bayes-based Context Extension for Large Language Models

Jianlin Su, Murtadha Ahmed, Wenbo, Luo Ao, Mingren Zhu, Yunfeng Liu

DeepMind

arXiv:2403.17552v115.733 citationsh-index: 56Has CodeNAACL

Originality Incremental advance

AI Analysis

This addresses a bottleneck for researchers and practitioners using LLMs by enabling more effective in-context learning with increased demonstrations, though it is an incremental improvement over existing methods.

The paper tackles the problem of transformer length limitations hindering in-context learning with many demonstrations by introducing NBCE, a framework that expands context size without fine-tuning, resulting in substantial performance gains, especially with more examples, and consistently outperforming alternatives.

Large Language Models (LLMs) have shown promising in-context learning abilities. However, conventional In-Context Learning (ICL) approaches are often impeded by length limitations of transformer architecture, which pose challenges when attempting to effectively integrate supervision from a substantial number of demonstration examples. In this paper, we introduce a novel framework, called Naive Bayes-based Context Extension (NBCE), to enable existing LLMs to perform ICL with an increased number of demonstrations by significantly expanding their context size. Importantly, this expansion does not require fine-tuning or dependence on particular model architectures, all the while preserving linear efficiency. NBCE initially splits the context into equal-sized windows fitting the target LLM's maximum length. Then, it introduces a voting mechanism to select the most relevant window, regarded as the posterior context. Finally, it employs Bayes' theorem to generate the test task. Our experimental results demonstrate that NBCE substantially enhances performance, particularly as the number of demonstration examples increases, consistently outperforming alternative methods. The NBCE code will be made publicly accessible. The code NBCE is available at: https://github.com/amurtadha/NBCE-master

View on arXiv PDF Code

Similar