0.6.0 - 2026-08-17
Summary
nvMolKit 0.6.0 adds two new GPU-accelerated capabilities: Maximum Common Substructure (MCS) search and a FIRE minimizer option for MMFF and UFF force field optimization. Fused Butina clustering has been rewritten on a pure CUDA backend, making it over 10x faster and removing the Triton dependency, and Butina clustering now supports reordering=False and matches RDKit output exactly. Substructure search is up to 2x faster end to end with a new DFS backend. This release supports RDKit 2025.09.1 through 2026.03.5 and includes all fixes from the v0.5.1 patch release.
Contributors
- Clay Moore (@mooreneural)
- Eva Xue (@evasnow1992)
- Kevin Boyd (@scal444)
- Matthew Neba (@Matthew-Neba)
- Timur Rvachov (@trvachov)
Features
- GPU-accelerated Maximum Common Substructure (MCS) search via the new
nvmolkit.mcs.findMCSAPI. Searches many molecule pairs in one batched call — all pairs of a molecule set, an explicit pair list, or two zipped lists — with configurable atom/bond matching, per-pair timeouts, and a CPU RDKit fallback for oversized inputs. RDKit fMCS options that are not yet supported raise informative errors (#221) - FIRE minimizer as an alternative to BFGS for force field optimization, roughly 2x faster than BFGS for comparable accuracy. Available via
minimizerKind="FIRE"in the MMFF and UFF optimization APIs and theBatchedForcefieldAPI (#216) - Fused Butina clustering rewritten from Python/Triton onto a CUDA backend, keeping the full clustering loop on the GPU. Up to 10x faster and removes the Triton dependency, by @Matthew-Neba (#226)
- Support for
reordering=Falsein Butina clustering, matching the RDKit option, by @Matthew-Neba (#192) - Butina clustering results now match RDKit exactly, with deterministic tie-breaking and cluster ordering, by @Matthew-Neba (#241)
- New DFS-based substructure search backend, up to 2x faster end to end, with reduced host-device transfer overhead (#244)
- Conformer RMSD can now return a square distance matrix (
output_format="square") for direct use with downstream clustering APIs (a26ae9c) - Performance improvements to conformer RMSD kernels by @mooreneural (ea09e2d)
- Reduced per-call launch overhead in similarity computations by @mooreneural (ea09e2d)
Bug Fixes
- All bug fixes from the v0.5.1 patch release, covering ETKDG conformer generation, Morgan fingerprints, fused Butina clustering, TFD, SMARTS handling, MMFF/UFF convergence, and pip packaging
- Fix a data race in batched MMFF minimization when a molecule's conformers spanned multiple batches, which could produce incorrect atom typing or crashes at high thread counts (#236)
- Fix MMFF failing to parametrize cyclophosphazine-type molecules (#258)
- Fix an incorrect sixth-order torsion gradient term in distance geometry minimization (#217)
- CUDA streams belonging to a different GPU are now rejected with a clear error instead of being silently accepted (#235)
Miscellaneous
- Document fixes for common torch CUDA-backend installation issues with pip, uv, and conda (#209)
- New downloadable agent skill teaching LLM coding agents to use nvMolKit, with evals and an NVIDIA skills catalog card (434440f)
- (C++) Major CMake / include-path refactor; all includes are now relative to the project root (#191)