6.2IRMar 27
Large language models for post-publication research evaluation: Evidence from expert recommendations and citation indicatorsMengjia Wu, Yi Zhang, Robin Haunschild et al.
Assessing the quality of scientific research is essential for scholarly communication, yet widely used approaches face limitations in scalability, subjectivity, and time delay. Recent advances in large language models (LLMs) offer new opportunities for automated research evaluation based on textual content. This study examines whether LLMs can support post-publication peer review tasks by benchmarking their outputs against expert judgments and citation-based indicators. Two evaluation tasks are constructed using articles from the H1 Connect platform: identifying high-quality articles and performing finer-grained evaluation including article rating, merit classification, and expert style commenting. Multiple model families, including BERT models, general-purpose LLMs, and reasoning oriented LLMs, are evaluated under multiple learning strategies. Results show that LLMs perform well in coarse grained evaluation tasks, achieving accuracy above 0.8 in identifying highly recommended articles. However, performance decreases substantially in fine-grained rating tasks. Few-shot prompting improves performance over zero-shot settings, while supervised fine-tuning produces the strongest and most balanced results. Retrieval augmented prompting improves classification accuracy in some cases but does not consistently strengthen alignment with citation indicators. The overall correlations between model outputs and citation indicators remain positive but moderate.
1.2DLDec 4, 2025
Introducing multiverse analysis to bibliometrics: The case of team size effects on disruptive researchChristian Leibel, Lutz Bornmann
Although bibliometrics has become an essential tool in the evaluation of research performance, bibliometric analyses are sensitive to a range of methodological choices. Subtle choices in data selection, indicator construction, and modeling decisions can substantially alter results. Ensuring robustness (meaning that findings hold up under different reasonable scenarios) is therefore critical for credible research and research evaluation. To address this issue, this study introduces multiverse analysis to bibliometrics. Multiverse analysis is a statistical tool that enables analysts to transparently discuss modeling assumptions and thoroughly assess model robustness. Whereas standard robustness checks usually cover only a small subset of all plausible models, multiverse analysis includes all plausible models. The benefits of multiverse analysis are illustrated by assessing the robustness of the findings reported by Wu et al. (2019), who observed that small teams tend to produce more disruptive research than large teams. While we found robust evidence of a negative effect of team size on disruption scores, the effect size depends substantially on the model specification. Our findings underscore the importance of assessing the multiverse robustness of bibliometric results to clarify their practical implications.
3.6IRJan 27, 2021
Investigating Dissemination of Scientific Information on Twitter: A Study of Topic Networks in Opioid PublicationsRobin Haunschild, Lutz Bornmann, Devendra Potnis et al.
One way to assess a certain aspect of the value of scientific research is to measure the attention it receives on social media. While previous research has mostly focused on the "number of mentions" of scientific research on social media, the current study applies "topic networks" to measure public attention to scientific research on Twitter. Topic networks are the networks of co-occurring author keywords in scholarly publications and networks of co-occurring hashtags in the tweets mentioning those scholarly publications. This study investigates which topics in opioid scholarly publications have received public attention on Twitter. Additionally, it investigates whether the topic networks generated from the publications tweeted by all accounts (bot and non-bot accounts) differ from those generated by non-bot accounts. Our analysis is based on a set of opioid scholarly publications from 2011 to 2019 and the tweets associated with them. We use co-occurrence network analysis to generate topic networks. Results indicated that Twitter users have mostly used generic terms to discuss opioid publications, such as "Opioid," "Pain," "Addiction," "Treatment," "Analgesics," "Abuse," "Overdose," and "Disorders." Results confirm that topic networks provide a legitimate method to visualize public discussions of health-related scholarly publications and how Twitter users discuss health-related scientific research differently from the scientific community. There was a substantial overlap between the topic networks based on the tweets by all accounts and non-bot accounts. This result indicates that it might not be necessary to exclude bot accounts for generating topic networks as they have a negligible impact on the results.
5.9DLAug 29, 2020
A Decade of In-text Citation Analysis based on Natural Language Processing and Machine Learning Techniques: An overview of empirical studiesSehrish Iqbal, Saeed-Ul Hassan, Naif Radi Aljohani et al.
Citation analysis is one of the most frequently used methods in research evaluation. We are seeing significant growth in citation analysis through bibliometric metadata, primarily due to the availability of citation databases such as the Web of Science, Scopus, Google Scholar, Microsoft Academic, and Dimensions. Due to better access to full-text publication corpora in recent years, information scientists have gone far beyond traditional bibliometrics by tapping into advancements in full-text data processing techniques to measure the impact of scientific publications in contextual terms. This has led to technical developments in citation context and content analysis, citation classifications, citation sentiment analysis, citation summarisation, and citation-based recommendation. This article aims to narratively review the studies on these developments. Its primary focus is on publications that have used natural language processing and machine learning techniques to analyse citations.