TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
For researchers developing large language models for biology, this work provides a unified, large-scale corpus and evaluation suite that significantly improves biological understanding without sacrificing general performance.
TheBioCollection is a 52.6B-token pre-training corpus for biology that unifies scattered biological data into a coherent format. Training a fixed 16B-parameter model on this corpus more than doubles its overall score on the accompanying evaluation suite, with gains across all biological domains while preserving general language ability.
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.