Retrieved article excerpt
Open article Β· Retrieved 2026-09-12T15:22:31.996157+00:00
# BioNeMo Inference Runtime
## Easy, fast, and memory-efficient structure prediction inference
GPU-accelerated inference for protein, nucleic-acid, and ligand structure
prediction models β from FASTA/MSA to PDB/mmCIF.
[Speedup against input size on H100](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/assets/speedup-vs-residues.png)
## About
BioNeMo Inference Runtime (BioIR) is NVIDIA's library for structure-prediction
inference. A five-stage GPU pipeline turns AlphaFold-lineage and all-atom models
into PDB/mmCIF with confidence scores. Models stay ordinary `nn.Module`s β no
TensorRT engine build.
## Getting Started
### Prerequisites
- **Linux, x86\_64 or aarch64**, with an NVIDIA GPU. The wheels are
`manylinux_2_34`, so the host needs glibc 2.34 or newer β Ubuntu 22.04, RHEL 9
or later.
- **Driver 580 or newer.** The dev image carries a CUDA 13.2 build of PyTorch.
An older driver runs it only through the forward-compatibility shim, which we
have measured hanging and crashing part-way through a run rather than merely
running slowly β results taken on one are discarded, not corrected.
- **Python 3.12.** The released wheels are tagged `cp312`, so pip finds no
matching build on a newer interpreter.
- **Docker** and the **NVIDIA Container Toolkit**, to build from source in the
dev container. Installing the wheel needs neither.
PyTorch and the CUDA math libraries arrive as wheel dependencies, or in
`nvcr.io/nvidia/pytorch:26.05-py3` when you use the container. Building the
extension from source outside a container needs a C++17 compiler and CUDA
headers as well β
[`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md#prerequisites-for-bioir-development-workflow).
### Release-qualified GPUs
H200, H100, A100, L40S, GB200 and GB300. Measured speedup, memory and accuracy
for each: [`docs/ref/benchmark.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md).
BioIR runs on more than these. The
[support matrix](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/support-matrix.md#gpus) lists every architecture the
backend covers and which fused kernels apply to each; those devices work but
are not part of this release's qualification.
### Install
BioIR is published on PyPI, one wheel per CPU architecture:
```
pip install bionemo-ir
```
The wheel ships the kernels precompiled, so nothing in the install builds CUDA
and running it needs only the driver's `libcuda.so.1`. That is the whole
install if you are calling BioIR from your own code β the container below is
for working on BioIR itself. [`docs/install.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/install.md) covers the
environment setup and the requirements in full.
### Build from source
Configure [SSH authentication with GitHub](https://docs.github.com/en/authentication/connecting-to-github-with-ssh), then clone the
repository and fetch its submodules and LFS objects.
```
git lfs install &&
GIT_LFS_SKIP_SMUDGE=0 \
git clone --recurse-submodules \
[email protected]:NVIDIA-BioNeMo/BioNeMo-Inference-Runtime.git &&
cd BioNeMo-Inference-Runtime
```
Then, build the dev image and open a shell in it:
```
docker/dev.sh
```
The image carries the dependencies; your checkout is bind-mounted, so install
the package once inside and fold something:
```
pip install -e '.[dev]'
scripts/fetch_weights.sh --model boltz-2
python examples/folding/run_demo.py --output-dir output
```
Checkpoints come from their upstream publishers and need no NVIDIA credentials;
anything that cannot be fetched is skipped, and the tests needing it skip too.
Running `scripts/run_tests.sh` stages weights and runs the suite the way CI
does.
Building without a container needs more than a Python environment β see
[`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md#prerequisites-for-bioir-development-workflow) for
the prerequisites and the wheel build. The rest of that page covers daily
development; [`docs/ref/docker-images.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/docker-images.md) covers the
images and what `docker/dev.sh` mounts.
## Documentation
BioIR documentation lives under [`docs/`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs) and is published with Fern:
- [Overview](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/fern/pages/overview.mdx) β what BioIR is and how to start
- [Installation](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/install.md) β requirements and release-wheel installation
- [Quickstart](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/quickstart.md) β run a serial Boltz-2 prediction
- [Ray multi-GPU inference](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ray.md) β scale independent requests across
visible GPUs
- [Developer guide](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md) β build, test, stage weights, contribute
- [API reference](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/api.md) β `build_processor`, model constructors,
inputs/outputs
- [Architecture](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/architecture.md) β five-stage pipeline and runtime
design
- [Config architecture](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/config.md) β model `BaseConfig` tree and
pipeline stage configs
- [Support matrix](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/support-matrix.md) β models, GPUs, and fused kernels
- [Benchmarks](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md) β measured speedup and memory against
OSS PyTorch
- [Model weights](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/model-weights.md) β checkpoint resolution and staging
- [Docker images](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/docker-images.md) β development and runtime images
- [Coding guidelines](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/coding.md) β style, naming, and tooling
- [Folding example](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/examples/folding/README.md) β runnable `build_processor`
demo
## Benchmarks
### Methodology
Folding benchmarks over a bench set the shipped
[`rebuild_dataset.py`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/.agents/skills/bench-perf-oss/dataset) builds from RCSB
and NVIDIA's MSA Search NIM β there is no dataset release to download.
Template-bearing samples included: both sides load every bundled MSA and attach
every listed template.
- One GPU, serial, one structure per forward call.
- Time only GPU-synchronized `model.forward()`. Featurization, transfers,
postprocessing, writing, and scoring stay outside the window.
- Discard one warmup forward, then report one measured forward per sample.
- BioIR runs its default optimized config, with a CUDA graph on the diffusion
module where supported.
- OSS runs its own inference script: eager always, plus `torch.compile` when it
passes a dynamic-shape probe.
- Runtime knobs match on both sides β 200 sampling steps, 3 or 5 diffusion
samples, and per-model recycling.
- Score written structures with OpenStructure lDDT and DockQ. Speedup is
`OSS forward / BioIR forward`; above 1 favors BioIR.
- Future work will add additional Blackwell-optimized kernels.
### Results
| Model | H100 | H200 |
| --- | --- | --- |
| Boltz-2 | 1.78x / 2.65x | 1.74x / 2.54x |
| OpenFold3 | 1.55x / 2.02x | 1.54x / 2.03x |
| OpenFold2 / AlphaFold2 monomer | 2.55x / 2.60x | 2.61x / 2.66x |
| OpenFold2 / AlphaFold2 multimer | 2.66x / 2.77x | 2.61x / 2.75x |
| Protenix | β / 1.87x | β / 1.84x |
Geomean speedup, `vs OSS torch.compile / vs OSS PyTorch eager`; above 1 favours
BioIR. Protenix has no `torch.compile` path. Fourteen GPUs, per-model accuracy
and peak memory, and how to reproduce any of it:
[`docs/ref/benchmark.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md).
The [`bench-perf-oss` agent skill](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/.agents/skills/bench-perf-oss/SKILL.md) has
the full gates, environment isolation, result schema, and charting protocol.
## Contributing
We welcome contributions. See [`contributing.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/contributing.md) for
policy and [`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md) for the development workflow.
## Citation
If you use BioIR in your research, please cite it via
[`CITATION.cff`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/CITATION.cff).
## Contact / Support
- Bugs and feature requests:
[GitHub Issues](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/issues/new/choose)
- Usage questions:
[GitHub Discussions](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/discussions)
- Security vulnerabilities: see [`SECURITY.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/SECURITY.md) β do **not**
file a public issue
## License
NVIDIA-authored BioIR code is licensed under the [Apache License 2.0](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/LICENSE).
Distribution compliance material is available here:
- [Third-party notices and attributions](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/THIRD_PARTY_NOTICES.md)
- [Full third-party license texts](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/LICENSES)
- [Gemmi 0.6.5 corresponding source](https://github.com/project-gemmi/gemmi/tree/v0.6.5), licensed under MPL-2.0 or
LGPL-3.0-or-later; BioIR distributes it under the MPL-2.0 option