2026-10-11 16:37 UTC

NVIDIA claims its released BioNeMo Inference Runtime accelerates supported structure-prediction model forward passes by roughly 1.5–2.7 times versus OSS torch.compile on H100 and H200 while retaining ordinary PyTorch modules, potentially lowering scientific inference costs without TensorRT engine builds.

state: seedheat: mediumuncertainty: mediumnovelscott: lowai-infrastructure inference-runtime inference-economics biomedical-aiNVIDIA

What is this?

NVIDIA has made BioNeMo Inference Runtime (BioIR) available as a Python library for biomolecular structure-prediction inference on NVIDIA GPUs; its announcement describes it as a public beta. It supplies specialized kernels and graph optimizations while keeping models as ordinary PyTorch nn.Modules, without TensorRT engine builds or export steps. NVIDIA's early benchmarks report geometric-mean forward-pass speedups of 1.54–2.61Γ— over OSS torch.compile for OpenFold3, Boltz2, and OpenFold2 monomer on H100 and H200. These are vendor-reported, GPU-synchronized model-forward measurements that exclude featurization, transfers, postprocessing, writing, and scoring, so the snippets do not establish equivalent end-to-end cost savings or independent validation.

Why it matters to Scott

Scott has tested TensorRT acceleration and uses PyTorch, but the hits establish no supported biomolecular workload or H100/H200 deployment that BioIR would improve; this is an adjacent runtime development, not a demonstrated extension or challenge to his positions. The radar tracks GPU kernels and inference economics, but no supplied page tracks BioIR itself, and NVIDIA’s forward-pass benchmarks do not establish end-to-end savings for Scott’s systems.
dev:technology.nvidia-tensorrtdev:technology.pytorchradar:concept.gpu-kernelsradar:concept.inference-economicsradar:concept.inference-engines
queries asked of Scott's wikis
  • inference economics end-to-end costs versus kernel benchmarks
  • PyTorch native inference versus compiled engine deployment
  • GPU optimization specialized kernels CUDA graphs
  • benchmark methodology baseline fairness reproducibility
  • biomolecular structure prediction scientific inference projects

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 697h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-12 15:27 (minted)⭐ origin echo-reconstructedNVIDIA publishes BioIR as an installable structure-prediction inference library with precompiled kernels, ordinary nn.Module models, and rep
NVIDIA on github (echo) Β· attributed from hn.story.49672745 Β· published time unknown
β€”
09-12 14:36first on hacker news Β· published Β· lag ?BioNeMo Inference Runtime
khoahpd
β€”
09-12 14:36amplified on hacker news πŸ‘‘hn.story.49672745
khoahpd
peak 4 Β· 0 comments Β· 98% of case engagement
09-12 15:20our radar first saw it Β· lag ?discovery anchor: hn.story.49672745β€”
pace: p23 vs 1032 stories at the 336h mark (now 697h old) β€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBioNeMo Inference Runtime
Retrieved article excerpt

Open article Β· Retrieved 2026-09-12T15:22:31.996157+00:00

# BioNeMo Inference Runtime

## Easy, fast, and memory-efficient structure prediction inference

GPU-accelerated inference for protein, nucleic-acid, and ligand structure
prediction models β€” from FASTA/MSA to PDB/mmCIF.

[Speedup against input size on H100](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/assets/speedup-vs-residues.png)

## About

BioNeMo Inference Runtime (BioIR) is NVIDIA's library for structure-prediction
inference. A five-stage GPU pipeline turns AlphaFold-lineage and all-atom models
into PDB/mmCIF with confidence scores. Models stay ordinary `nn.Module`s β€” no
TensorRT engine build.

## Getting Started

### Prerequisites

- **Linux, x86\_64 or aarch64**, with an NVIDIA GPU. The wheels are
  `manylinux_2_34`, so the host needs glibc 2.34 or newer β€” Ubuntu 22.04, RHEL 9
  or later.
- **Driver 580 or newer.** The dev image carries a CUDA 13.2 build of PyTorch.
  An older driver runs it only through the forward-compatibility shim, which we
  have measured hanging and crashing part-way through a run rather than merely
  running slowly β€” results taken on one are discarded, not corrected.
- **Python 3.12.** The released wheels are tagged `cp312`, so pip finds no
  matching build on a newer interpreter.
- **Docker** and the **NVIDIA Container Toolkit**, to build from source in the
  dev container. Installing the wheel needs neither.

PyTorch and the CUDA math libraries arrive as wheel dependencies, or in
`nvcr.io/nvidia/pytorch:26.05-py3` when you use the container. Building the
extension from source outside a container needs a C++17 compiler and CUDA
headers as well β€”
[`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md#prerequisites-for-bioir-development-workflow).

### Release-qualified GPUs

H200, H100, A100, L40S, GB200 and GB300. Measured speedup, memory and accuracy
for each: [`docs/ref/benchmark.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md).

BioIR runs on more than these. The
[support matrix](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/support-matrix.md#gpus) lists every architecture the
backend covers and which fused kernels apply to each; those devices work but
are not part of this release's qualification.

### Install

BioIR is published on PyPI, one wheel per CPU architecture:

```
pip install bionemo-ir
```

The wheel ships the kernels precompiled, so nothing in the install builds CUDA
and running it needs only the driver's `libcuda.so.1`. That is the whole
install if you are calling BioIR from your own code β€” the container below is
for working on BioIR itself. [`docs/install.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/install.md) covers the
environment setup and the requirements in full.

### Build from source

Configure [SSH authentication with GitHub](https://docs.github.com/en/authentication/connecting-to-github-with-ssh), then clone the
repository and fetch its submodules and LFS objects.

```
git lfs install &&
  GIT_LFS_SKIP_SMUDGE=0 \
    git clone --recurse-submodules \
      [email protected]:NVIDIA-BioNeMo/BioNeMo-Inference-Runtime.git &&
  cd BioNeMo-Inference-Runtime
```

Then, build the dev image and open a shell in it:

```
docker/dev.sh
```

The image carries the dependencies; your checkout is bind-mounted, so install
the package once inside and fold something:

```
pip install -e '.[dev]'
scripts/fetch_weights.sh --model boltz-2
python examples/folding/run_demo.py --output-dir output
```

Checkpoints come from their upstream publishers and need no NVIDIA credentials;
anything that cannot be fetched is skipped, and the tests needing it skip too.
Running `scripts/run_tests.sh` stages weights and runs the suite the way CI
does.

Building without a container needs more than a Python environment β€” see
[`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md#prerequisites-for-bioir-development-workflow) for
the prerequisites and the wheel build. The rest of that page covers daily
development; [`docs/ref/docker-images.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/docker-images.md) covers the
images and what `docker/dev.sh` mounts.

## Documentation

BioIR documentation lives under [`docs/`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs) and is published with Fern:

- [Overview](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/fern/pages/overview.mdx) β€” what BioIR is and how to start
- [Installation](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/install.md) β€” requirements and release-wheel installation
- [Quickstart](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/quickstart.md) β€” run a serial Boltz-2 prediction
- [Ray multi-GPU inference](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ray.md) β€” scale independent requests across
  visible GPUs
- [Developer guide](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md) β€” build, test, stage weights, contribute
- [API reference](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/api.md) β€” `build_processor`, model constructors,
  inputs/outputs
- [Architecture](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/architecture.md) β€” five-stage pipeline and runtime
  design
- [Config architecture](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/config.md) β€” model `BaseConfig` tree and
  pipeline stage configs
- [Support matrix](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/support-matrix.md) β€” models, GPUs, and fused kernels
- [Benchmarks](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md) β€” measured speedup and memory against
  OSS PyTorch
- [Model weights](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/model-weights.md) β€” checkpoint resolution and staging
- [Docker images](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/docker-images.md) β€” development and runtime images
- [Coding guidelines](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/coding.md) β€” style, naming, and tooling
- [Folding example](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/examples/folding/README.md) β€” runnable `build_processor`
  demo

## Benchmarks

### Methodology

Folding benchmarks over a bench set the shipped
[`rebuild_dataset.py`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/.agents/skills/bench-perf-oss/dataset) builds from RCSB
and NVIDIA's MSA Search NIM β€” there is no dataset release to download.
Template-bearing samples included: both sides load every bundled MSA and attach
every listed template.

- One GPU, serial, one structure per forward call.
- Time only GPU-synchronized `model.forward()`. Featurization, transfers,
  postprocessing, writing, and scoring stay outside the window.
- Discard one warmup forward, then report one measured forward per sample.
- BioIR runs its default optimized config, with a CUDA graph on the diffusion
  module where supported.
- OSS runs its own inference script: eager always, plus `torch.compile` when it
  passes a dynamic-shape probe.
- Runtime knobs match on both sides β€” 200 sampling steps, 3 or 5 diffusion
  samples, and per-model recycling.
- Score written structures with OpenStructure lDDT and DockQ. Speedup is
  `OSS forward / BioIR forward`; above 1 favors BioIR.
- Future work will add additional Blackwell-optimized kernels.

### Results

| Model | H100 | H200 |
| --- | --- | --- |
| Boltz-2 | 1.78x / 2.65x | 1.74x / 2.54x |
| OpenFold3 | 1.55x / 2.02x | 1.54x / 2.03x |
| OpenFold2 / AlphaFold2 monomer | 2.55x / 2.60x | 2.61x / 2.66x |
| OpenFold2 / AlphaFold2 multimer | 2.66x / 2.77x | 2.61x / 2.75x |
| Protenix | β€” / 1.87x | β€” / 1.84x |

Geomean speedup, `vs OSS torch.compile / vs OSS PyTorch eager`; above 1 favours
BioIR. Protenix has no `torch.compile` path. Fourteen GPUs, per-model accuracy
and peak memory, and how to reproduce any of it:
[`docs/ref/benchmark.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/ref/benchmark.md).

The [`bench-perf-oss` agent skill](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/.agents/skills/bench-perf-oss/SKILL.md) has
the full gates, environment isolation, result schema, and charting protocol.

## Contributing

We welcome contributions. See [`contributing.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/contributing.md) for
policy and [`docs/dev.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/dev.md) for the development workflow.

## Citation

If you use BioIR in your research, please cite it via
[`CITATION.cff`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/CITATION.cff).

## Contact / Support

- Bugs and feature requests:
  [GitHub Issues](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/issues/new/choose)
- Usage questions:
  [GitHub Discussions](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/discussions)
- Security vulnerabilities: see [`SECURITY.md`](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/docs/SECURITY.md) β€” do **not**
  file a public issue

## License

NVIDIA-authored BioIR code is licensed under the [Apache License 2.0](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/LICENSE).
Distribution compliance material is available here:

- [Third-party notices and attributions](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/THIRD_PARTY_NOTICES.md)
- [Full third-party license texts](https://github.com/NVIDIA-BioNeMo/BioNeMo-Inference-Runtime/blob/main/LICENSES)
- [Gemmi 0.6.5 corresponding source](https://github.com/project-gemmi/gemmi/tree/v0.6.5), licensed under MPL-2.0 or
  LGPL-3.0-or-later; BioIR distributes it under the MPL-2.0 option
khoahpd40
🟧 echo.github ⭐NVIDIA publishes BioIR as an installable structure-prediction inference library with precompiled kernels, ordinary nn.Module models, and repNVIDIAβ€”β€”

Interpretation history

Decision trace