Research Entity Extraction and Topic Detection from UKRI Grant Proposals
For funding agencies and policymakers, this work provides a practical evaluation of LLM-based methods for analyzing sensitive grant data to identify emerging research areas.
This paper compares three LLM-based approaches for extracting and classifying research entities from UKRI grant proposals, finding that Mistral and GPT-4o outperform a bespoke algorithm, with Mistral achieving 90.5% topic classification accuracy versus 71.4% for the DSIT-Taxonomies pipeline.
This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke algorithm, DSIT-Taxonomies, for extracting and classifying research entities from funding proposals. Our project "Tracking Stars and Unicorns" aims to identify early signals of emerging research areas to inform public investment. Our methodology employed a three-stage pipeline, leveraging Mistral for primary entity extraction and mapping against the OpenAlex Topics taxonomy. We evaluated our approach across 42 proposals' abstracts from different areas and observed that Mistral and GPT-4o produce comparable, high-quality entity sets with significant semantic overlap, outperforming the fragmented DSIT-Taxonomies approach. Crucially, the Mistral-based approach achieved superior topic classification accuracy (90.5%) compared to the full DSIT-Taxonomies pipeline (71.4%). We conclude that Mistral offers a high-performance, operationally efficient, and secure solution for large-scale analysis of sensitive grant data.