CL LGJan 7, 2025

Multilingual Open QA on the MIA Shared Task

Navya Yarrabelly, Saloni Mittal, Ketan Todi, Kimihiro Hasegawa

arXiv:2501.04153v14.91 citationsh-index: 3

Originality Incremental advance

AI Analysis

This work addresses the challenge of accessing information across languages for users in low-resource settings, though it is incremental as it builds on existing retrieval methods.

The paper tackles the problem of cross-lingual information retrieval for open question answering in low-resource languages by proposing a zero-shot re-ranking method using a multilingual question generation model, achieving improved passage retrieval without requiring labeled data.

Cross-lingual information retrieval (CLIR) ~\cite{shi2021cross, asai2021one, jiang2020cross} for example, can find relevant text in any language such as English(high resource) or Telugu (low resource) even when the query is posed in a different, possibly low-resource, language. In this work, we aim to develop useful CLIR models for this constrained, yet important, setting where we do not require any kind of additional supervision or labelled data for retrieval task and hence can work effectively for low-resource languages. \par We propose a simple and effective re-ranking method for improving passage retrieval in open question answering. The re-ranker re-scores retrieved passages with a zero-shot multilingual question generation model, which is a pre-trained language model, to compute the probability of the input question in the target language conditioned on a retrieved passage, which can be possibly in a different language. We evaluate our method in a completely zero shot setting and doesn't require any training. Thus the main advantage of our method is that our approach can be used to re-rank results obtained by any sparse retrieval methods like BM-25. This eliminates the need for obtaining expensive labelled corpus required for the retrieval tasks and hence can be used for low resource languages.

View on arXiv PDF

Similar