Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
For SSH researchers, this work addresses the challenge of adapting LLMs to disciplinary diversity and multilingual sources, though it is an ongoing use case with no final results.
The paper presents a use case within the LLMs4EU project to adapt foundation models for SSH research tasks like question answering and literature review, with evaluation including quantitative benchmarks and qualitative expert assessment.
The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially with regard to disciplinary diversity, multilingual access to sources and the evaluation of results. This paper presents an on-going use case developed within the European project LLMs4EU and the ALT-EDIC infrastructure, aimed at adapting foundation models to SSH research practices and supporting tasks such as question answering, comparative document analysis and literature review. The evaluation framework follows the LLMs4EU protocol and encompasses both independent quantitative benchmarking (retrieval, summarisation, traceability and hallucination detection) and a qualitative assessment involving a panel of Digital Humanities experts. By embedding model adaptation within research infrastructures and a structured legal and ethical compliance framework, the use case explores how domain-sensitive and regulation-aware generative AI can support SSH scholarship while preserving reliability and epistemic responsibility.