LGAug 8, 2024

NFDI4Health workflow and service for synthetic data generation, assessment and risk management

arXiv:2408.04478v12 citationsh-index: 8
Originality Synthesis-oriented
AI Analysis

This work provides a practical solution for researchers needing synthetic health data to advance AI while protecting patient privacy, though it is incremental in integrating existing tools into a service framework.

The paper presents a workflow and services for generating, assessing, and managing synthetic health data to address privacy concerns in AI development, showcasing tools like VAMBN, MultiNODEs, and SYNDAT using datasets from ADNI and RKI.

Individual health data is crucial for scientific advancements, particularly in developing Artificial Intelligence (AI); however, sharing real patient information is often restricted due to privacy concerns. A promising solution to this challenge is synthetic data generation. This technique creates entirely new datasets that mimic the statistical properties of real data, while preserving confidential patient information. In this paper, we present the workflow and different services developed in the context of Germany's National Data Infrastructure project NFDI4Health. First, two state-of-the-art AI tools (namely, VAMBN and MultiNODEs) for generating synthetic health data are outlined. Further, we introduce SYNDAT (a public web-based tool) which allows users to visualize and assess the quality and risk of synthetic data provided by desired generative models. Additionally, the utility of the proposed methods and the web-based tool is showcased using data from Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Center for Cancer Registry Data of the Robert Koch Institute (RKI).

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes