BMAILGOct 25, 2024

Multi-view biomedical foundation models for molecule-target and property prediction

IBM
arXiv:2410.19704v42.36 citationsh-index: 28Has CodeAdv Sci
Originality Incremental advance
AI Analysis

This work addresses the need for more reliable molecular predictions in drug discovery, though it is incremental as it builds on existing single-view foundation models.

The authors tackled the problem of creating robust molecular representations for biomedical tasks by developing MMELON, a multi-view foundation model that integrates graph, image, and text views, which matched the performance of the best single-view model across over 120 tasks including molecular solubility and GPCR activity.

Quality molecular representations are key to foundation model development in bio-medical research. Previous efforts have typically focused on a single representation or molecular view, which may have strengths or weaknesses on a given task. We develop Multi-view Molecular Embedding with Late Fusion (MMELON), an approach that integrates graph, image and text views in a foundation model setting and may be readily extended to additional representations. Single-view foundation models are each pre-trained on a dataset of up to 200M molecules. The multi-view model performs robustly, matching the performance of the highest-ranked single-view. It is validated on over 120 tasks, including molecular solubility, ADME properties, and activity against G Protein-Coupled receptors (GPCRs). We identify 33 GPCRs that are related to Alzheimer's disease and employ the multi-view model to select strong binders from a compound screen. Predictions are validated through structure-based modeling and identification of key binding motifs.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes