CLLGFeb 6, 2023

Efficient and Flexible Topic Modeling using Pretrained Embeddings and Bag of Sentences

arXiv:2302.03106v34 citationsh-index: 12Has Code
Originality Incremental advance
AI Analysis

This work addresses the challenge of improving topic modeling for NLP applications by making it more efficient and flexible, though it is incremental in leveraging existing embeddings.

The paper tackles the problem of topic modeling by proposing a novel algorithm that uses a bag of sentences approach with pre-trained sentence embeddings, achieving state-of-the-art results with relatively low computational demands.

Pre-trained language models have led to a new state-of-the-art in many NLP tasks. However, for topic modeling, statistical generative models such as LDA are still prevalent, which do not easily allow incorporating contextual word vectors. They might yield topics that do not align well with human judgment. In this work, we propose a novel topic modeling and inference algorithm. We suggest a bag of sentences (BoS) approach using sentences as the unit of analysis. We leverage pre-trained sentence embeddings by combining generative process models and clustering. We derive a fast inference algorithm based on expectation maximization, hard assignments, and an annealing process. The evaluation shows that our method yields state-of-the art results with relatively little computational demands. Our method is also more flexible compared to prior works leveraging word embeddings, since it provides the possibility to customize topic-document distributions using priors. Code and data is at \url{https://github.com/JohnTailor/BertSenClu}.

Code Implementations2 repos
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes