Draft:BioEmu
| BioEmu | |
|---|---|
| Developer | Microsoft Research AI for Science |
| Release | February 20, 2025 |
| Written in | Python |
| Operating system | Linux |
| Type | Protein structure prediction and molecular modelling |
| License | MIT License |
| Website | github |
BioEmu (short for Biomolecular Emulator) is an open-source generative deep learning model developed by Microsoft Research's AI for Science group. Given a amino acid sequence, it generates protein-backbone conformations intended to approximate the equilibrium distribution of structures adopted by a protein monomer.[1] Unlike systems aimed mainly at predicting one representative structure, such as AlphaFold, BioEmu is intended to represent multiple conformational states and their relative probabilities.[2][3] Its outputs are independent samples, not a time-ordered molecular dynamics trajectory, and therefore do not by themselves give transition paths or rates.[2]
Development and release
[edit]A research team led by Frank Noé developed BioEmu at Microsoft Research AI for Science.[2][1] The project was intended to reduce the computational cost of studying protein conformational ensembles, for which conventional molecular-dynamics simulations may require long runs to observe rare structural changes.[4]
A preprint describing BioEmu was posted in December 2024. Microsoft released the inference code and model weights on 20 February 2025. The peer-reviewed paper was published online on 10 July 2025 and appeared in the 14 August issue of Science.[5][4][1][2] The software and weights are distributed under the MIT License.[6]
Method
[edit]BioEmu uses sequence representations derived with AlphaFold2 and ColabFold as input to a diffusion model based on the Distributional Graphormer (DiG) architecture.[7][6] Its training combined static structures from the AlphaFold Protein Structure Database, more than 200 milliseconds of molecular-dynamics simulations, and experimental measurements of protein stability.[1] The model can generate thousands of statistically independent structures per hour on one graphics processing unit.[1] It produces backbone frames; side chains can be reconstructed and the resulting structures can be relaxed with a short molecular-dynamics calculation.[6]
In benchmarks reported by its developers, BioEmu covered 85% of reference domain motions. The reported mean absolute error for metastable-state free energies was 0.91 kcal/mol relative to long molecular-dynamics simulations, while the error for folding free energies was 0.76 kcal/mol relative to experimental measurements.[7]
Limitations
[edit]BioEmu is designed mainly for monomeric soluble proteins under fixed conditions. It does not explicitly account for ligands, membranes, oligomerisation, pH or temperature, and its developers report low prediction quality for multichain protein interactions.[7][8] Because part of its training data consists of AlphaFold2 predictions and molecular-dynamics simulations, errors and biases in those sources may be inherited by the model.[7]
Applications
[edit]Haitin, Manori and Giladi used BioEmu to generate conformational pools for small-angle X-ray scattering analysis. Pools of about 1,000 conformers were sufficient to fit 20 deposited datasets after selection against the experimental data, with fit quality comparable to the published analyses. The unweighted BioEmu ensembles did not reproduce the scattering data, so the authors treated them as priors rather than direct estimates of equilibrium populations.[8]
Bhakat and Strauch combined BioEmu structures with short molecular-dynamics simulations and Markov state models for the kinases BRAF and CDK2. Their workflow sampled functionally relevant metastable states and the population shift caused by the BRAF V600E mutation more effectively than an AlphaFold2 reduced-MSA approach. It did not reproduce experimentally observed heterogeneity in GlyT1 or plasmepsin II, where side-chain changes were important.[9]
Dusza, Wodziński and Szaleniec used BioEmu to generate starting structures for atomistic simulations of a tungsten-containing aldehyde oxidoreductase. They reconstructed cofactors for its three subunits and minimized the resulting holo models; 241 of the 287 conformations returned by BioEmu passed quality control. The retained structures differed from the cryo-electron microscopy reference, while their cofactor geometry remained close to that of control models derived from the experimental structure.[10]
In a virtual-screening benchmark, Shin, Joo and Yoo generated nearly 1,300 BioEmu structures across 26 kinases. The ensembles were structurally diverse, but consensus and best-score docking did not improve on crystal-structure baselines and often performed worse. The authors concluded that selecting useful conformations, rather than generating diversity, was the main bottleneck for this use.[11]
See also
[edit]External links
[edit]References
[edit]- 1 2 3 4 5 Lewis, Sarah; Hempel, Tim; Jiménez-Luna, José; et al. (August 14, 2025). "Scalable emulation of protein equilibrium ensembles with generative deep learning". Science. 389 (6761) eadv9817. doi:10.1126/science.adv9817. ISSN 0036-8075. PMID 40638710.
- 1 2 3 4 Barnhart, Max (10 July 2025). "Microsoft AI predicts protein conformations". Chemical & Engineering News. Retrieved 13 September 2026.
- ↑ Singh, Arunima (October 2025). "BioEmu is a biomolecular emulator for sampling protein structure ensembles". Nature Methods. 22 (10): 2008. doi:10.1038/s41592-025-02874-1. ISSN 1548-7091. PMID 41068462.
- 1 2 Lewis, Sarah; Hempel, Tim; Jiménez-Luna, José (20 February 2025). "Exploring the structural changes driving protein function with BioEmu-1". Microsoft Research. Retrieved 13 September 2026.
- ↑ Lewis, Sarah; Hempel, Tim; Jiménez Luna, José; et al. (December 2024). "Scalable emulation of protein equilibrium ensembles with generative deep learning". doi.org. doi:10.1101/2024.12.05.626885. Retrieved 2026-09-24.
- 1 2 3 "microsoft/bioemu". GitHub. Retrieved 13 September 2026.
- 1 2 3 4 "BioEmu model card". Hugging Face. Microsoft. Retrieved 13 September 2026.
- 1 2 Haitin, Yoni; Manori, Bar; Giladi, Moshe (June 2026). "Use of biomolecular emulator for characterizing flexible proteins by small-angle x-ray scattering". Protein Science. 35 (6) e70646. doi:10.1002/pro.70646. ISSN 0961-8368. PMC 13240575. PMID 42178622.
- ↑ Bhakat, Soumendranath; Strauch, Eva-Maria (June 2026). "Accelerated Sampling of Protein Dynamics Using BioEmu-Augmented Molecular Simulation". Journal of Chemical Information and Modeling. 66 (12): 7168–7178. doi:10.1021/acs.jcim.6c01000. ISSN 1549-9596. PMID 42268730.
- ↑ Dusza, Piotr; Wodziński, Marek; Szaleniec, Maciej (2027). Gervasi, Osvaldo; Murgante, Beniamino; Garau, Chiara; Rocha, Ana Maria A. C. (eds.). From BioEmu Monomer Ensembles to AMBER-Ready Holo Models: Subunit-Aware Quality Control of a Cofactor-Rich Aldehyde Oxidoreductase. Computational Science and Its Applications – ICCSA 2026 Workshops. Lecture Notes in Computer Science. Vol. 16765. Cham: Springer Nature Switzerland. pp. 193–211. doi:10.1007/978-3-032-30536-7_13. ISBN 978-3-032-30535-0. Retrieved 2026-09-24.
- ↑ Shin, Jaeoh; Joo, Keehyoung; Yoo, Jejoong (2026). "Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck". Journal of Chemical Information and Modeling. 66 (16): 9905–9915. doi:10.1021/acs.jcim.6c01677. ISSN 1549-9596. PMID 42670848.