LGCEQMJun 29

Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

arXiv:2607.162637.5h-index: 18
Predicted impact top 42% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For antibody design researchers, this provides a scalable solution to optimize antibody expressibility in data-constrained settings, though the improvement is incremental over existing methods.

The paper tackles antibody expression ranking, a critical task in antibody design hindered by scarce labeled data. By integrating scarce quantitative data with large-scale weak positive supervision from immunization data via a preference-based learning framework, the method consistently outperforms baselines on most metrics, evaluated on 1254 labeled and 4 million unlabeled sequences.

Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. We adapt Direct Preference Optimization (DPO) to protein language models by introducing a union-masked log-likelihood approximation and IMGT-based alignment, enabling efficient training on variable-length sequences. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid-derived antibodies, we show that our method consistently outperforms baselines on most metrics. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data-constrained settings. Project page: https://kisoji-biotechnology-inc.github.io/Preference-Expression-Ranking/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes