Skip to content

v0.6.0

Latest

Choose a tag to compare

@scal444 scal444 released this 17 Aug 17:45

0.6.0 - 2026-08-17

Summary

nvMolKit 0.6.0 adds two new GPU-accelerated capabilities: Maximum Common Substructure (MCS) search and a FIRE minimizer option for MMFF and UFF force field optimization. Fused Butina clustering has been rewritten on a pure CUDA backend, making it over 10x faster and removing the Triton dependency, and Butina clustering now supports reordering=False and matches RDKit output exactly. Substructure search is up to 2x faster end to end with a new DFS backend. This release supports RDKit 2025.09.1 through 2026.03.5 and includes all fixes from the v0.5.1 patch release.

Contributors

Features

  • GPU-accelerated Maximum Common Substructure (MCS) search via the new nvmolkit.mcs.findMCS API. Searches many molecule pairs in one batched call — all pairs of a molecule set, an explicit pair list, or two zipped lists — with configurable atom/bond matching, per-pair timeouts, and a CPU RDKit fallback for oversized inputs. RDKit fMCS options that are not yet supported raise informative errors (#221)
  • FIRE minimizer as an alternative to BFGS for force field optimization, roughly 2x faster than BFGS for comparable accuracy. Available via minimizerKind="FIRE" in the MMFF and UFF optimization APIs and the BatchedForcefield API (#216)
  • Fused Butina clustering rewritten from Python/Triton onto a CUDA backend, keeping the full clustering loop on the GPU. Up to 10x faster and removes the Triton dependency, by @Matthew-Neba (#226)
  • Support for reordering=False in Butina clustering, matching the RDKit option, by @Matthew-Neba (#192)
  • Butina clustering results now match RDKit exactly, with deterministic tie-breaking and cluster ordering, by @Matthew-Neba (#241)
  • New DFS-based substructure search backend, up to 2x faster end to end, with reduced host-device transfer overhead (#244)
  • Conformer RMSD can now return a square distance matrix (output_format="square") for direct use with downstream clustering APIs (a26ae9c)
  • Performance improvements to conformer RMSD kernels by @mooreneural (ea09e2d)
  • Reduced per-call launch overhead in similarity computations by @mooreneural (ea09e2d)

Bug Fixes

  • All bug fixes from the v0.5.1 patch release, covering ETKDG conformer generation, Morgan fingerprints, fused Butina clustering, TFD, SMARTS handling, MMFF/UFF convergence, and pip packaging
  • Fix a data race in batched MMFF minimization when a molecule's conformers spanned multiple batches, which could produce incorrect atom typing or crashes at high thread counts (#236)
  • Fix MMFF failing to parametrize cyclophosphazine-type molecules (#258)
  • Fix an incorrect sixth-order torsion gradient term in distance geometry minimization (#217)
  • CUDA streams belonging to a different GPU are now rejected with a clear error instead of being silently accepted (#235)

Miscellaneous

  • Document fixes for common torch CUDA-backend installation issues with pip, uv, and conda (#209)
  • New downloadable agent skill teaching LLM coding agents to use nvMolKit, with evals and an NVIDIA skills catalog card (434440f)
  • (C++) Major CMake / include-path refactor; all includes are now relative to the project root (#191)