Skip to content

Unvendor the bundled C++ deps (protobuf, abseil, onnx, re2, flatbuffers, coremltools), and add an optional TensorRT EP - #183

Open
hmaarrfk wants to merge 13 commits into
conda-forge:mainfrom
hmaarrfk:coreml-libcoremltools
Open

hmaarrfk wants to merge 13 commits into
conda-forge:mainfrom
hmaarrfk:coreml-libcoremltools

Conversation

@hmaarrfk

@hmaarrfk hmaarrfk commented May 17, 2026 •

Copy link
Copy Markdown
Contributor

I've been wanting to help unvendor lots of the stuff that this repo depends on.

The list is:

I'm not really sure how valuable this is, but it is in the spirit of conda-forge as a whole.

While we were in here we also picked up the TensorRT EP that #199 has been asking
for. It turns out to be pretty small as long as CUDA stays in tree, so it rides
along as an opt-in onnxruntime-ep-tensorrt package instead of a new variant
axis, and the job count stays where it is.

61 of the 71 jobs are green. The 10 that are not are the CUDA 13.4 linux ones,
and they only fail because the new output name is not registered yet
(conda-forge/admin-requests#2362). They build and pass their tests before they
get to that gate.

So i figured I would ask my AI tools to do the hardest parts of my OSS contributions.

  • I have reviewed AI's code
  • This code is ready to be read by other humans.
AI's details

What this does

onnxruntime 1.30.0 FetchContent-builds and statically links its own copies of
many dependencies. This branch makes it link the conda-forge shared libraries
instead:

Dependency Change
libprotobuf host dep; find_package already supported. Needs onnxruntime_USE_FULL_PROTOBUF=ON and --path_to_protoc_exe pointing at the build-prefix protoc, so the generated code matches the runtime ABI
libabseil host dep + patch 0011 (onnxruntime pins an exact-version FIND_PACKAGE_ARGS; abseil's CMake config is ExactVersion, so the constraint must be dropped)
libcoremltools patch 0010 — the CoreML EP links the prebuilt library instead of vendoring the coremltools sources (osx-arm64)
libonnx host dep, no onnxruntime patch needed any more — the required symbols are exported by libonnx itself
re2 host dep; find_package already supported — no patch
flatbuffers build + host dep; build.sh/bld.bat regenerate the schema headers with the conda flatc (the committed *.fbs.h carry a static_assert pinning the flatbuffers version)

Patch 0012 switches four callers from protobuf's deprecated
RepeatedField::Resize to resize(), because Windows builds with /sdl and
that turns the resulting C4996 into an error.

Patch 0013 backports the source hunks of microsoft/onnxruntime#32359, still
open upstream, so the recipe compiles against ONNX 1.23.0. 1.23 adds the
FLOAT6E2M3 and FLOAT6E3M2 TensorProto data types and opset 28; onnxruntime
1.30.0 was released against 1.22, so its element-size table is sized by
TensorProto_DataType_DataType_ARRAYSIZE, the two new slots default to zero,
and it trips its own static_assert. Three upstream hunks are deliberately
omitted, with reasons in the patch header.

Upstream work this depended on (all merged)

Linking a shared onnx needed real changes on the onnx side:

Two onnxruntime-feedstock fixes were split out of this branch and merged
separately: #209 (the Windows novec package was silently vectorized) and #211
(CUDA 13.4).

The TensorRT execution provider

#199 flagged that the plugin direction (#197/#198/#203) leaves TensorRT with
nothing to build against, since upstream has no TRT EP plugin. This branch
keeps CUDA in tree, which is the case #199 itself called "the classic in-tree
approach", so none of that applies and neither of #204's onnxruntime patches is
needed. #200 established the cmake flags.

onnxruntime_USE_TENSORRT=ON with the builtin parser, so the shared
libnvonnxparser is linked rather than onnx-tensorrt vendored. Enabled by the
presence of the TensorRT headers in the host prefix, so meta.yaml alone decides
which variants get it. Scope is CUDA 13.4 on linux, decided by availability
rather than preference: conda-forge's TensorRT 11.1 needs __glibc >=2.28 and
libstdcxx >=15 for its CUDA 13 build, which matches that variant and not the
CUDA 12.9 one (glibc 2.17).

libnvinfer is 1.9 GB, so it must not reach anyone who did not ask for it:

  • onnxruntime_providers_tensorrt is a standalone module library and
    libonnxruntime never links it, so the core packages gain the provider-bridge
    entry points but no link dependency. Those run_exports are dropped with
    ignore_run_exports_from.
  • upstream's setup.py bundles every provider library it finds into the wheel's
    capi/ directory, so a TensorRT-enabled build drops
    libonnxruntime_providers_tensorrt.so in there. build.sh removes it,
    otherwise the core package would ship a shared object nothing can load.
  • a new onnxruntime-ep-tensorrt output installs that library into
    site-packages/onnxruntime/capi, where onnxruntime actually looks for an
    in-tree provider (Env::GetRuntimePath() is the directory of the loaded
    libonnxruntime). That is what makes
    providers=["TensorrtExecutionProvider"] work with no registration call.

It builds inside the existing CUDA 13.4 jobs, so the matrix stays at 71. Its
dependency on onnxruntime is an exact build pin, not a range: the provider
bridge is an unversioned internal C++ ABI, a ProviderHost vtable handed to
Provider_SetHost with no negotiation and no error on mismatch.

Testing

Reviews raised the concern that moving to shared libraries can hide failures
until runtime. Everything that can go wrong there traces to one thing — onnx and
onnxruntime now share one copy of the onnx protobuf descriptors and one ONNX
schema registry — and it all shows up in one scenario: build a model with onnx,
have onnxruntime run it, then use onnx again. So there is one flat test of that
per codepath rather than a suite:

  • recipe/test_onnxruntime_runtime.py (69 lines) imports onnxruntime before onnx
    (the order that actually broke), builds and checks and shape-infers with onnx,
    runs at ORT_ENABLE_ALL so the optimizers register contrib schemas into the
    shared registry, compares against onnx's reference evaluator on every provider
    the hardware allows, then re-checks the model with onnx.
  • recipe/cpp_interop/ (103 lines) does the same through the onnx C++ API and a
    second Ort::Env. This is what catches a missing export or a duplicate
    descriptor registration in the onnxruntime-cpp output.
  • recipe/test_ep_tensorrt.py (51 lines) loads the EP library with RTLD_NOW
    from the directory onnxruntime searches, which catches both ways a bad install
    fails — wrong directory, or unresolved libnvinfer — without needing a GPU.
    With a GPU it runs a model instead.

onnxruntime's own C++ unit tests also run now: build.sh/bld.bat point
onnx_SOURCE_DIR at the .proto files shipped by the onnx package, which
re-enables them under onnxruntime_USE_FULL_PROTOBUF=ON, and the recipe runs
them through ctest. osx builds run natively on the GitHub Actions arm runners.

Verified on an RTX 4090 against locally built CUDA 13.4 packages, not only in
CI: TensorRT runs a model with every node placed on the EP (read back from the
profile) and matches the reference evaluator; the core package still works with
the EP package absent; and the EP test fails loudly in an environment where the
library is missing.

Still vendored

cpuinfo (pytorch/cpuinfo) and kleidiai remain vendored and statically
linked — there is no conda-forge package for either yet.

Known limitations

  • The conda-forge libonnx does not carry onnxruntime's vendored patch that
    un-deprecates the opset-18 GroupNormalization operator, so models with an
    ai.onnx GroupNormalization node at opset 18–20 are rejected. opset ≥ 21 is
    fine, and no mainstream exporter emits that operator. The corresponding
    onnxruntime unit tests are filtered out on osx.
  • recipe/patch_absl_injected_base_names.py is a temporary workaround: on
    Windows with CUDA 13.4, nvcc's host pass rejects abseil's
    friend class MixingHashState::HashStateBase; injected-class-name
    declarations. nvcc 13.0 accepts the same headers and onnxruntime's own
    vendored abseil has identical code, so this is an nvcc regression. It should
    move upstream rather than live here.
  • get_available_providers() lists TensorrtExecutionProvider even when
    onnxruntime-ep-tensorrt is not installed, because the provider-bridge entry
    points are compiled into the core. Asking for it anyway logs an error and
    falls back to CPU rather than crashing, but the message is upstream's own
    "Please install TensorRT libraries … make sure they're in the PATH or
    LD_LIBRARY_PATH", which is misleading here — the answer is
    conda install onnxruntime-ep-tensorrt.
  • onnxruntime-ep-tensorrt covers linux-64 and linux-aarch64. win-64 has the
    conda-forge TensorRT packages and would be a follow-up; the C++
    onnxruntime-cpp consumer is not covered either, since that would need the
    library next to $PREFIX/lib/libonnxruntime.so as well.
  • onnxruntime-ep-tensorrt is a new output name and needs registering in
    conda-forge/feedstock-outputs (Add onnxruntime-ep-tensorrt to the onnxruntime feedstock admin-requests#2362). Until that
    lands the CUDA 13.4 linux jobs go red on output validation after building and
    testing successfully. onnxruntime-ep-cuda from Split recipe: CPU packages + standalone CUDA EP plugin (70 CI jobs -> 14) #203 is unregistered too.

https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe

@conda-forge-admin

conda-forge-admin commented May 17, 2026 •

Copy link
Copy Markdown
Contributor

Hi! This is the friendly automated conda-forge-linting service.

I just wanted to let you know that I linted all conda-recipes in your PR (recipe/meta.yaml) and found it was in an excellent condition.

I do have some suggestions for making it better though...

For recipe/meta.yaml:

  • ℹ️ The recipe is not parsable by parser conda-souschef (grayskull). This parser is not currently used by conda-forge, but may be in the future. We are collecting information to see which recipes are compatible with grayskull.
  • ℹ️ The recipe is not parsable by parser conda-recipe-manager. The recipe can only be automatically migrated to the new v1 format if it is parseable by conda-recipe-manager.

This message was generated by GitHub Actions workflow run https://github.com/conda-forge/conda-forge-webservices/actions/runs/35948423481. Examine the logs at this URL for more detail.

Comment thread recipe/build.sh Outdated
Comment on lines +13 to +20
# The C++ unit tests build an onnx_test_data_proto target that imports ONNX's
# .proto source files. The unvendored conda-forge libonnx package ships only
# the generated headers, not the .proto sources, so the unit tests cannot be
# built. Skip them; the recipe's own test section still exercises the Python
# package, the C++ consumer test and cmake-package-check.
echo "Compiled unit tests are disabled"
RUN_TESTS_BUILD_PY_OPTIONS=""
BUILD_UNIT_TESTS="OFF"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this might be the biggest change. im pretty ok with this but i should triple check that the proto files arent shipped somewhere

Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/meta.yaml Outdated
Comment thread recipe/build.sh Outdated
Comment thread recipe/build.sh Outdated
hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

The OpenVINO EP pulls libopenvino-onnx-frontend into the host prefix, which
hard-depends on the conda-forge global libprotobuf (6.x). onnxruntime's
vendored onnx must be generated and compiled against that same protobuf, or its
headers get shadowed by the 6.x ones in $PREFIX/include and onnx fails to
compile ("onnx::ModelProto is incomplete type"). So on linux-64:

- build: use the global libprotobuf (provides a 6.x protoc) instead of pinning
  3.21; keep 3.21 on other platforms (unchanged, proven).
- host: add libprotobuf so the vendored onnx is compiled against the same 6.x
  headers/protoc as libopenvino-dev. onnxruntime then find_package()s the conda
  protobuf and links it instead of vendoring 3.21.
- license_file: protobuf and its abseil dependency are no longer vendored on
  linux-64 (they come from the conda libprotobuf/libabseil packages), so their
  _deps source trees don't exist there -- gate those license entries to
  [not linux64].

onnxruntime 1.26 / onnx 1.21 compile cleanly under protobuf 6.x: the only
removed API (RepeatedPtrField::ReleaseCleared) is guarded for >=5.26, and onnx
1.21 already requires abseil-based protobuf >=4.25.1. This is the same direction
as the draft unvendoring PR conda-forge#183.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

The OpenVINO EP pulls libopenvino-onnx-frontend into the host prefix, which
hard-pins the conda libprotobuf it was built against (6.33). onnxruntime's
build protoc and vendored onnx must use that same protobuf, or the 6.x host
headers shadow the vendored 3.21 ones and onnx fails to compile
("onnx::ModelProto is incomplete type"). So on linux-64:

- build + host: pin libprotobuf 6.33.* to match OpenVINO 2026.1. The conda-forge
  global pin has already moved to 7.x, so an unpinned libprotobuf gives a 7.x
  protoc in build vs OpenVINO's 6.33 in host -- another skew; pin both to 6.33
  (bump when OpenVINO migrates to 7.x). Other platforms keep the proven 3.21.
- onnxruntime then find_package()s the conda protobuf and links it instead of
  vendoring 3.21.
- license_file: protobuf is no longer vendored on linux-64 (it comes from the
  conda libprotobuf package), so gate the protobuf-src license to [not linux64].

onnxruntime 1.26 / onnx 1.21 compile cleanly under protobuf 6.x: the only
removed API (RepeatedPtrField::ReleaseCleared) is guarded for >=5.26, and onnx
1.21 already requires abseil-based protobuf >=4.25.1. Same direction as the
draft unvendoring PR conda-forge#183.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
@hmaarrfk
hmaarrfk marked this pull request as ready for review June 2, 2026 00:48
@hmaarrfk

hmaarrfk commented Jun 2, 2026

Copy link
Copy Markdown
Contributor Author

One huge advantage of this is that you can start to have openvino as a runtime (which requires libprotobuf)

hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

Builds on the unvendored-dependencies work (PR conda-forge#183), so onnxruntime already
links the conda-forge libprotobuf 6.33.5 / libabseil 20260107 that OpenVINO
2026.1 also uses -- no protobuf version skew.

build.sh: pass --use_openvino CPU on linux-64. The "CPU" argument is only the
default device; the resulting libonnxruntime_providers_openvino.so is a
separately loadable module and the device is chosen at runtime via
provider_options=[{"device_type": "CPU"|"GPU"|"NPU"|"AUTO:GPU,CPU"|...}].
onnxruntime finds OpenVINO via find_package(OpenVINO REQUIRED COMPONENTS
Runtime ONNX), satisfied by libopenvino-dev on CMAKE_PREFIX_PATH.

meta.yaml (all gated to linux64):
- host: libopenvino-dev for the OpenVINO CMake config + onnx frontend.
- OpenVINO stays a *soft* dependency: ignore_run_exports_from libopenvino-dev /
  libopenvino / libopenvino-onnx-frontend / libopenvino-ir-frontend (the only
  OpenVINO packages with run_exports), and express compatibility via
  run_constrained: pin_compatible('libopenvino'). Every OpenVINO frontend and
  device plugin depends on an exact libopenvino build, so constraining
  libopenvino alone transitively pins them all. OpenVINO is therefore installed
  only when the user asks for it, and the OpenVINO EP picks it up when present.
- test: assert OpenVINOExecutionProvider is advertised (in a bare env without
  OpenVINO, since the provider library is loaded lazily).

migrations: add libopenvino_dev20261 to pin libopenvino 2026.1.0 (the build the
EP is compiled against). The OpenVINO EP is Linux x86_64 only; mirrors how the
CoreML EP is baked into the default osx-arm64 build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

Apply the libopenvino_dev20261 migration to the linux-64 variant configs so the
OpenVINO EP is built against (and run_constrained to) libopenvino 2026.1.0,
which uses the same libprotobuf 6.33.5 / libabseil 20260107 onnxruntime already
links via PR conda-forge#183's absl_grpc_proto_26Q1 migration.

Only the .ci_support/*.yaml libopenvino_dev pins are included; the conda-smithy
scaffolding churn from a full rerender (this machine's conda-smithy is newer than
PR conda-forge#183's) is intentionally left out so this stacks cleanly on PR conda-forge#183. A normal
rerender at merge time regenerates the rest.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
@hmaarrfk

hmaarrfk commented Jun 2, 2026

Copy link
Copy Markdown
Contributor Author

i think the biggest question is whether or not we want to try to unvendor onnx with libonnx. i think it will be worthwhile.

Comments indicating approval on conda-forge/onnx-feedstock#146 would be appreciated.

If we want to keep libonnx vendored, then i can make those changes too.

hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

Builds on the unvendored-dependencies work (PR conda-forge#183), so onnxruntime already
links the conda-forge libprotobuf 6.33.5 / libabseil 20260107 that OpenVINO
2026.1 also uses -- no protobuf version skew.

build.sh: pass --use_openvino CPU on linux-64. The "CPU" argument is only the
default device; the resulting libonnxruntime_providers_openvino.so is a
separately loadable module and the device is chosen at runtime via
provider_options=[{"device_type": "CPU"|"GPU"|"NPU"|"AUTO:GPU,CPU"|...}].
onnxruntime finds OpenVINO via find_package(OpenVINO REQUIRED COMPONENTS
Runtime ONNX), satisfied by libopenvino-dev on CMAKE_PREFIX_PATH.

meta.yaml (all gated to linux64):
- host: libopenvino-dev for the OpenVINO CMake config + onnx frontend.
- OpenVINO stays a *soft* dependency: ignore_run_exports_from libopenvino-dev /
  libopenvino / libopenvino-onnx-frontend / libopenvino-ir-frontend (the only
  OpenVINO packages with run_exports), and express compatibility via
  run_constrained: pin_compatible('libopenvino'). Every OpenVINO frontend and
  device plugin depends on an exact libopenvino build, so constraining
  libopenvino alone transitively pins them all. OpenVINO is therefore installed
  only when the user asks for it, and the OpenVINO EP picks it up when present.
- test: assert OpenVINOExecutionProvider is advertised (in a bare env without
  OpenVINO, since the provider library is loaded lazily).

migrations: add libopenvino_dev20261 to pin libopenvino 2026.1.0 (the build the
EP is compiled against). The OpenVINO EP is Linux x86_64 only; mirrors how the
CoreML EP is baked into the default osx-arm64 build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
hmaarrfk added a commit to hmaarrfk/onnxruntime-feedstock that referenced this pull request Jun 2, 2026
<details><summary>Claude's draft</summary>

Apply the libopenvino_dev20261 migration to the linux-64 variant configs so the
OpenVINO EP is built against (and run_constrained to) libopenvino 2026.1.0,
which uses the same libprotobuf 6.33.5 / libabseil 20260107 onnxruntime already
links via PR conda-forge#183's absl_grpc_proto_26Q1 migration.

Only the .ci_support/*.yaml libopenvino_dev pins are included; the conda-smithy
scaffolding churn from a full rerender (this machine's conda-smithy is newer than
PR conda-forge#183's) is intentionally left out so this stacks cleanly on PR conda-forge#183. A normal
rerender at merge time regenerates the rest.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Resume this Claude session:
```
claude --resume 8959e104-fade-4e06-a903-ec84f609fa1c
```
</details>
@cbourjau

cbourjau commented Jun 2, 2026

Copy link
Copy Markdown
Contributor

i think the biggest question is whether or not we want to try to unvendor onnx with libonnx. i think it will be worthwhile.

Comments indicating approval on conda-forge/onnx-feedstock#146 would be appreciated.

If we want to keep libonnx vendored, then i can make those changes too.

I gave a small update concerning libonnx in the linked PR. Could/should we unvendor onnx/protobuf and the rest in two separate PRs? There was also work upstream to make the execution providers separate from the main package (e.g. the CUDA kernels don't have to be statically linked to libonnxruntime anymore). I think this would a great thing to make use of in this feedstock. Is it in scope of this exercise, too?

@hmaarrfk

hmaarrfk commented Jun 2, 2026

Copy link
Copy Markdown
Contributor Author

to understand:

  • One PR for unvendoring libonnx
  • One PR for unvendoring protobuf and the rest

is that correct?

hmaarrfk added a commit to hmaarrfk/onnx-feedstock that referenced this pull request Sep 14, 2026
<details><summary>Claude's draft</summary>

A runtime test for conda-forge/onnxruntime-feedstock#183 found that with
the shared libonnx, importing onnxruntime and then onnx.reference fails:

  NotImplementedError: This function assumes every operator has a unique
  name 'Unique' even across multiple domains 'com.microsoft' and ''.

The ONNX schema registry now lives in libonnx and is shared by everything
in the process, so onnxruntime's com.microsoft schemas show up in
onnx.defs.get_all_schemas_with_history(). onnx.reference indexes every
schema by name and refuses clashes across domains. Before unvendoring,
onnxruntime had its own private registry. Upstream main has the same code.

Patch 0009 skips schemas outside ONNX's own domains in _build_schemas,
since those are the only operators the reference evaluator implements.
test_reference_foreign_domain.py reproduces it without onnxruntime by
registering a Relu in another domain: it fails with the same error on
onnx build 3 and passes with the patch.

Verified locally on linux-64: all four onnx builds pass the new test. An
onnxruntime linux-64 build ran its new runtime test suite with the
patched onnx.reference, and everything passes.

Resume this Claude session:
```
cd /home/mark/git/feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
conda-forge-admin pushed a commit to conda-forge/onnx-feedstock that referenced this pull request Sep 14, 2026
<details><summary>Claude's draft</summary>

With libonnx build 3, onnxruntime on win-64 is down from 154 unresolved
externals to three (conda-forge/onnxruntime-feedstock#183):

  onnxruntime_graph.lib(contrib_defs.cc.obj) : error LNK2019: unresolved
    external symbol "__declspec(dllimport) ... onnx::ParseData<__int64>(...)"
  ... onnx::ParseData<int>(...), onnx::ToTensor<__int64>(...)

ToTensor and ParseData are declared as ONNX_API templates but their
explicit specializations are defined in tensor_proto_util.cc and
tensor_util.cc. GCC and clang give an explicit specialization the
visibility of its primary template, which is why linux and osx link.
MSVC only exports a specialization whose definition carries dllexport,
and build 3's onnx.dll exports none of them. Annotate the definitions,
and the ParseData(const Tensor*) declaration, with ONNX_API in patch
0008. Those are the only explicit specializations in the ONNX library
sources (cpp2py_export.cc's belong to the Python module).

Verified locally on linux-64: libonnx and onnx build and the specialization
files compile without warnings. win-64 is verified by this PR's CI only.

Resume this Claude session:
```
cd /home/mark/git/feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
conda-forge-admin pushed a commit to conda-forge/onnx-feedstock that referenced this pull request Sep 14, 2026
<details><summary>Claude's draft</summary>

onnxruntime_autoep_test aborted on osx-arm64 in
conda-forge/onnxruntime-feedstock#183 while loading its libraries:

  E descriptor_database.cc:683] File already exists in database: onnx/onnx-ml.proto
  F descriptor.cc:2531] Check failed: GeneratedDatabase()->Add(encoded_file_descriptor, size)

This is the "same protobuf definitions loaded twice" failure raised when
onnxruntime's protobuf was first unvendored. onnx's CMake compiles the
onnx_proto objects into the onnx library and also builds onnx_proto as a
library of its own, so libonnx and libonnx_proto both contain
onnx-ml.pb.cc (llvm-nm shows descriptor_table_onnx_2fonnx_2dml_2eproto in
both osx-arm64 dylibs, and libonnx does not link libonnx_proto). On
Linux, symbol interposition merges the two copies. macOS two-level
namespaces keep them apart, so any process loading both dylibs aborts.

New patch 0010: for shared builds, compile the protobuf sources in an
OBJECT library that only onnx pulls in, and keep onnx_proto as an
INTERFACE target linking onnx, so ONNX::onnx_proto still works for
consumers (onnxruntime lists both). Static builds are unchanged.
libonnx_proto.so / onnx_proto.dll are no longer shipped; the package
tests now assert they are absent.

Verified locally on linux-64: all five outputs build and pass their tests.
The new package contains only libonnx.so, which defines the descriptor
table once; ONNX::onnx_proto is INTERFACE_LINK_LIBRARIES ONNX::onnx; the
libonnx exception test links through the interface target; and
onnx_cpp2py_export needs only libonnx.so. osx and win: CI only.

Resume this Claude session:
```
cd /home/mark/git/feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
@hmaarrfk
hmaarrfk force-pushed the coreml-libcoremltools branch from 2c08ad5 to 77a5022 Compare September 17, 2026 18:33
@hmaarrfk hmaarrfk mentioned this pull request Sep 19, 2026
@hmaarrfk
hmaarrfk force-pushed the coreml-libcoremltools branch 2 times, most recently from bbbc631 to e8de978 Compare September 20, 2026 13:41
@hmaarrfk

Copy link
Copy Markdown
Contributor Author

@cbourjau @xhochy i would like to revive this conversation

basically there are 2 strong integration tests, on C++ test and one python test.

CoreML today is the best way onnxruntime has to run models on an apple processor, and if we are ok with the spirit to add "more integration tests" and merge in the spirit of unvendoring.

PS. ONNX 1.23 definitely threw a wrench in things (some of you predicted this), so i had to add a draft PR as a patch from upstream to help with compatibility. Can we "ignore" this hiccup for now?

@hmaarrfk
hmaarrfk force-pushed the coreml-libcoremltools branch from 063d6f0 to f0c22ac Compare September 22, 2026 19:13
<details><summary>Claude's draft</summary>

onnxruntime FetchContent-builds and statically links its own copies of
these libraries. Link the conda-forge shared libraries instead, in the
spirit of the rest of conda-forge.

  * libprotobuf, re2: find_package is already supported upstream, so
    they only need to be host dependencies. protobuf additionally needs
    onnxruntime_USE_FULL_PROTOBUF=ON (the conda build is full protobuf,
    not lite) and --path_to_protoc_exe pointing at the build-prefix
    protoc so the generated code matches the runtime ABI.
  * libabseil: patch 0011. onnxruntime pins an exact version through
    FIND_PACKAGE_ARGS and abseil's CMake config is ExactVersion, so the
    constraint has to be dropped for any other abseil to be accepted.
  * libonnx: a host dependency; no onnxruntime patch is needed because
    the symbols it links are now exported by libonnx itself
    (conda-forge/onnx-feedstock#157, conda-forge#158 and conda-forge#159).
  * flatbuffers: the committed *.fbs.h carry a static_assert pinning the
    flatbuffers version, so the build scripts regenerate them with the
    conda flatc. flatc must run on the build platform, hence the
    build-section dependency alongside the host one.

The C++ unit tests need ONNX's .proto sources, which libonnx does not
ship. Rather than re-vendoring onnx, point onnx_SOURCE_DIR at the .proto
files in the python onnx package and run the suite directly through
ctest. build.py's own python test phase stays off: it runs ONNX
conformance and quantization tooling tests that are brittle against an
external onnx and expect vendored test data.

Two GTEST_FILTER exclusions: the flaky QDQ MatMulNBits determinism
assertion (reads uninitialized memory upstream), and on osx the
GroupNormalization tests, which only exist for CoreML and depend on
onnxruntime's vendored patch un-deprecating the opset-18 operator --
conda-forge's libonnx follows upstream ONNX and rejects it. That
limitation is documented in meta.yaml.

The onnxruntime-cpp output repackages the prebuilt libraries rather than
compiling them, so conda-build cannot infer what libonnxruntime links.
List the libraries in its host section so their run_exports reach it,
otherwise downstream C++ consumers fail to link.

Drop the license files of the no-longer-vendored dependencies.

Windows needed the same treatment in bld.bat, plus three problems that
do not arise on unix.

CMAKE_DISABLE_FIND_PACKAGE_Protobuf=ON existed precisely to keep
conda-forge's protobuf out, and goes away.

libprotobuf is a DLL, and its CMake target defines PROTOBUF_USE_DLLS only
on targets that link it. onnxruntime_flatbuffers includes the onnx .pb.h
headers without linking protobuf, emits protobuf inline functions and
collides with the DLL's exports (LNK2005). Define it for every
translation unit. It has to go through cl.exe's CL variable rather than
CXXFLAGS, because build.py passes -DCMAKE_CXX_FLAGS whenever --parallel
is used, which makes CMake ignore CXXFLAGS -- that is why the first
attempt worked for CPU and failed for CUDA.

Windows builds with /sdl, which promotes C4996 to an error, and protobuf
has deprecated RepeatedField::Resize. Patch 0012 moves the four callers
to resize().

Locating onnx's .proto sources on Windows cannot rely on %SP_DIR% alone:
conda-build hardcodes Lib\site-packages, but conda-forge's python 3.15
moved to lib\python\site-packages (cfep-27). conda/conda-build#6143 fixes
it upstream but is not in any release yet (26.7.1), so probe both layouts
and fail loudly if neither has the file.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

On osx-arm64 the CoreML execution provider vendors the coremltools
sources and builds them as part of onnxruntime. Patch 0010 makes it link
the packaged libcoremltools instead, and adds the library as a host
dependency of both outputs.

coremltools was the only BSD-3-Clause component, so the osx-arm64 license
expression collapses back to the same "MIT AND BSL-1.0" used everywhere
else, and its license file is no longer collected. fp16 and psimd are
still vendored, so theirs stay.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

The win-64 CUDA 13.4 configs fail to compile abseil's hash headers. nvcc
13.4's host pass rejects declarations that name a base class through the
derived class's injected-class-name, e.g.

    friend class MixingHashState::HashStateBase;

nvcc 13.0 accepts the same headers, and onnxruntime's own vendored abseil
contains identical code, so this is an nvcc regression rather than
anything specific to the conda-forge package -- see
NVIDIA/cutlass#3065 for the same shape of
problem.

Four compiler-flag theories were tried first and none worked:
/Zc:twoPhase, /permissive-, /Zc:__cplusplus- and a partial header
rewrite. What does work is naming the bases explicitly at all six sites.
In the flat_hash_map/flat_hash_set/node_hash_map/node_hash_set headers
that also means renaming the Hash and Eq template parameters, because
raw_hash_map and raw_hash_set declare private `using Hash` / `using Eq`
members that hide them.

patch_absl_injected_base_names.py does this rewrite in place against the
host and build prefixes before the build. It is idempotent and exits
non-zero if any expected declaration is missing, so a future abseil that
has changed these lines fails loudly instead of silently not being
patched.

This belongs upstream, as a libabseil patch or an NVIDIA bug report, and
should be deleted once either lands.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

conda-forge's global pin moved libonnx to 1.23.0. onnxruntime 1.30.0 was
released against ONNX 1.22 and does not compile against it:

    onnxruntime/core/framework/tensorprotoutils.h:76:58: error:
      static assertion failed: TensorProto data type element-size map
      must cover every numeric data type.

ONNX 1.23.0 adds the TensorProto data types FLOAT6E2M3 (27) and
FLOAT6E3M2 (28) and bumps the ai.onnx opset to 28. onnxruntime's
element-size table is sized by TensorProto_DataType_DataType_ARRAYSIZE,
so the two new slots default to zero and trip its own static_assert;
onnxruntime_map_type_info.h has a matching consteval bijectivity check
that fails next.

Patch 0013 backports the source hunks of
microsoft/onnxruntime#32359 ("Integrate ONNX
1.23.0 base plumbing"), which is still open upstream: the two new enum
values and their mappings both ways, the element-size entries, opset
27 -> 28 in the transpose optimizer and the op builders (Q/DQ builders
stay at 27, as upstream), the EmbedLayerNorm fusion opset lists, and
session.allow_released_opsets_only on the in-memory load path. The C++
unit-test hunks are included so the ctest suite still builds.

Three upstream hunks are omitted, with the reasons recorded in the patch
header: the webgpu Logger refactor (unrelated, does not apply to 1.30.0,
and that provider is not built here), a qdq_transformer test that does
not exist in 1.30.0, and the cmake/vcpkg/requirements pins for
onnxruntime's own vendored ONNX, which this branch replaces with the
packaged libonnx.

Pin onnx to 1.23.0: builds 0 and 1 of onnx 1.22.0 predate the libonnx
split, ship their own libonnx.so and libonnx_proto.so and do not depend
on libonnx, so they would install alongside it and overwrite its files.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

Review of this branch raised the concern that moving from statically
linked private copies to shared libraries can hide failures until
runtime: duplicate protobuf descriptor or ONNX schema registration when
onnx and onnxruntime share a process, symbols missing from libonnx,
flatbuffers headers that no longer match the runtime, or execution
providers that fail to load. None of that is visible to an import check.

test_onnxruntime_runtime.py exercises it, comparing against onnx's pure
python ReferenceEvaluator so the assertions are about numerical results
rather than "it did not crash":

  * both import orders in fresh interpreters, plus onnx.reference
    imported after onnxruntime;
  * six models -- MLP, CNN, an ONNX function op, a model-local function,
    ai.onnx.ml ops and a transformer block -- at ORT_DISABLE_ALL and
    ORT_ENABLE_ALL, so the graph optimizers actually run;
  * a com.microsoft contrib fusion plus a protobuf round trip;
  * the ORT flatbuffers format and a LoRA adapter, which are what break
    if the flatbuffers schema headers drift;
  * CoreML on osx-arm64 with CPU fallback disabled, asserting through
    the profile that the nodes really landed on the EP;
  * the CUDA provider library, skipped when no NVIDIA driver is present
    (libcuda.so.1 comes from the driver, never from conda), and the full
    model suite on the GPU when ONNXRUNTIME_TEST_REQUIRE_CUDA=1.

cpp_interop/ covers the onnxruntime-cpp output the same way from C++: it
builds a ModelProto with the onnx C++ API, shape-infers it, runs it
through onnxruntime twice with separate Ort::Env instances and re-checks
it. This is what catches a missing export or a duplicate descriptor
registration.

The CoreML provider availability one-liner is subsumed by the CoreML
check in the new script.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

osx-arm64 was built on the default provider. GitHub Actions now offers
native Apple Silicon runners, so point the provider there and drop the
build_platform entry that was pinning osx_arm64 to itself.

This matters beyond build time: the CoreML checks in the runtime tests
only mean something on real Apple Silicon.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

conda-forge#211 published build 1 of onnxruntime 1.30.0, so this branch needs 2 to
supersede it.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
Other tools:
- conda-build 26.7.1
- rattler-build 0.76.0
- rattler-build-conda-compat 1.4.19
@hmaarrfk
hmaarrfk force-pushed the coreml-libcoremltools branch from f0c22ac to 6603b3f Compare September 22, 2026 20:56
<details><summary>Claude's draft</summary>

Addresses conda-forge#199, which asked for the TensorRT EP and flagged that the
plugin direction (conda-forge#197/conda-forge#198/conda-forge#203) has nothing for TensorRT to build
against upstream. This branch keeps CUDA in-tree, which is the case conda-forge#199
itself called "the classic in-tree approach and should be feasible", so
none of that applies here and neither of conda-forge#204's two onnxruntime patches
is needed. conda-forge#200 established the cmake flags; this is the same approach on
the v0 recipe, packaged differently.

Build side: onnxruntime_USE_TENSORRT=ON with the builtin parser, so the
shared libnvonnxparser is linked rather than onnx-tensorrt vendored --
consistent with the rest of this branch. The flags are gated on the
TensorRT headers being present in the host prefix, so meta.yaml alone
decides which variants get them.

Scope is CUDA 13.4 on linux, decided by availability rather than by
preference: conda-forge's TensorRT 11.1 needs __glibc >=2.28 and
libstdcxx >=15 for its CUDA 13 build, which matches that variant (glibc
2.28, gcc 15) and not the CUDA 12.9 one (glibc 2.17).

Packaging side, the part that took some care. libnvinfer is 1.9 GB, so it
must not reach anyone who did not ask for TensorRT:

  * onnxruntime_providers_tensorrt is a standalone module library and
    libonnxruntime never links it, so the core packages gain the
    provider-bridge entry points but no link dependency. Their
    run_exports are dropped with ignore_run_exports_from.
  * upstream's setup.py bundles every provider library it finds into the
    wheel's capi/ directory, so a TensorRT-enabled build drops
    libonnxruntime_providers_tensorrt.so in there, NEEDing libnvinfer.
    build.sh removes it, otherwise the core package would ship a shared
    object nothing can load.
  * a new onnxruntime-ep-tensorrt output installs that library into
    site-packages/onnxruntime/capi, where onnxruntime actually looks for
    an in-tree provider (Env::GetRuntimePath() is the directory of the
    loaded libonnxruntime, which for the python package is capi/). That
    is what makes providers=["TensorrtExecutionProvider"] work with no
    registration call and no patch.

It is built inside the existing CUDA 13.4 jobs, so the matrix stays at 71
jobs rather than growing a tensorrt variant axis.

The run dependency on onnxruntime is an exact build pin, not a range: the
provider bridge is an unversioned internal C++ ABI -- a ProviderHost
vtable handed to Provider_SetHost, with no negotiation and no error on
mismatch -- so a range would let a mismatched pair load and then
misbehave.

test_ep_tensorrt.py checks what an import cannot: that the library is in
the directory onnxruntime will search, that its NEEDED entries resolve
(a missing libnvinfer fails here rather than opaquely inside a later
session), that the provider is compiled into the core, and -- when a
driver and GPU are present -- that a model really runs through TensorRT,
compared against onnx's reference evaluator.

Not covered yet: C++ consumers of onnxruntime-cpp, which would need the
library next to $PREFIX/lib/libonnxruntime.so as well; win-64, which has
the conda-forge TensorRT packages; and linux-aarch64, which builds here
but cannot be run without hardware.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

onnxruntime_providers_tensorrt does not only link TensorRT.
cmake/onnxruntime_providers_tensorrt.cmake links it against onnx,
PROTOBUF_LIB, ABSEIL_LIBS, CUDA::cudart and onnxruntime_providers_shared
as well, and on this branch the first three are the shared conda-forge
libraries rather than vendored static ones. The new output's host section
had only the TensorRT packages, so conda-build's shared-object check would
have failed to resolve those NEEDED entries, and the run_exports that pin
them would never have reached the package.

Add libonnx, libprotobuf, libabseil and cuda-cudart-dev to its host.

onnxruntime_providers_shared is the one entry that cannot be resolved that
way: it is not a conda package but a library built here and shipped inside
the python onnxruntime package's capi/ directory. This output installs
into that same directory, so $ORIGIN resolves it at load time, and the
exact pin on onnxruntime guarantees the matching copy is present. It is
whitelisted for that reason and nothing else is --
test_ep_tensorrt.py loads it explicitly before the EP, so a broken
assumption here fails a test rather than passing silently.

Found by reading the upstream cmake rather than by CI: the local build
that would have caught it was killed for memory. Two related non-problems
confirmed at the same time, both worth recording because they would each
have needed a separate fix:

  * the core onnxruntime.cmake has no nvinfer or tensorrt reference at
    all, so libonnxruntime really does not link TensorRT and dropping
    those run_exports from the core packages is safe;
  * the provider is installed without an EXPORT set, so the cmake package
    config that onnxruntime-cpp ships does not gain a target pointing at
    a library that is not in it, and cmake-package-check is unaffected.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
@conda-forge-admin

Copy link
Copy Markdown
Contributor

Hi! This is the friendly automated conda-forge-linting service.

I wanted to let you know that I linted all conda-recipes in your PR (recipe/meta.yaml) and found some lint.

Here's what I've got...

For recipe/meta.yaml:

  • ❌ The recipe is not parsable by any of the known recipe parsers (['conda-forge-tick (the bot)', 'conda-recipe-manager', 'conda-souschef (grayskull)']). Please check the logs for more information and ensure your recipe can be parsed.

For recipe/meta.yaml:

  • ℹ️ The recipe is not parsable by parser conda-forge-tick (the bot). Your recipe may not receive automatic updates and/or may not be compatible with conda-forge's infrastructure. Please check the logs for more information and ensure your recipe can be parsed.
  • ℹ️ The recipe is not parsable by parser conda-souschef (grayskull). This parser is not currently used by conda-forge, but may be in the future. We are collecting information to see which recipes are compatible with grayskull.
  • ℹ️ The recipe is not parsable by parser conda-recipe-manager. The recipe can only be automatically migrated to the new v1 format if it is parseable by conda-recipe-manager.

This message was generated by GitHub Actions workflow run https://github.com/conda-forge/conda-forge-webservices/actions/runs/35950759554. Examine the logs at this URL for more detail.

<details><summary>Claude's draft</summary>

The previous commit added missing_dso_whitelist in a second `build:`
mapping, next to the one that already held `skip` and `string`. YAML keeps
the last duplicate key, so `skip` and `string` were silently discarded and
the TensorRT output was no longer gated to CUDA 13.4 linux. It then tried
to build on every platform -- CPU linux, aarch64, osx-arm64 -- where
libnvinfer does not exist and
build-ci/Release/libonnxruntime_providers_tensorrt.so was never built. 15
jobs failed, and so did the linter.

Merge the two mappings into one.

The gating was verified when the output was first added and then broken by
a later edit without re-checking it, which is exactly the failure this
kind of change invites. Re-verified now across five variants: the output
is present on linux CUDA 13.4 and absent on CPU linux, CUDA 12.9,
osx-arm64 and win-64 including win CUDA 13.4. Also checked that `string`,
`missing_dso_whitelist`, `script`, host and run all survive together, and
scanned the whole file for duplicate mapping keys -- the only remaining
ones are the selector-guarded `skip` and `script: install-cpp.sh/.bat`
pairs, which conda-build resolves before parsing.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
@conda-forge-admin

conda-forge-admin commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Hi! This is the friendly automated conda-forge-linting service.

I just wanted to let you know that I linted all conda-recipes in your PR (recipe/meta.yaml) and found it was in an excellent condition.

I do have some suggestions for making it better though...

For recipe/meta.yaml:

  • ℹ️ The recipe is not parsable by parser conda-souschef (grayskull). This parser is not currently used by conda-forge, but may be in the future. We are collecting information to see which recipes are compatible with grayskull.
  • ℹ️ The recipe is not parsable by parser conda-recipe-manager. The recipe can only be automatically migrated to the new v1 format if it is parseable by conda-recipe-manager.

This message was generated by GitHub Actions workflow run https://github.com/conda-forge/conda-forge-webservices/actions/runs/36079894167. Examine the logs at this URL for more detail.

<details><summary>Claude's draft</summary>

libonnx and libabseil were added to the new output's host on the strength
of cmake/onnxruntime_providers_tensorrt.cmake listing onnx and ABSEIL_LIBS
among the target's link libraries. CI shows that reading was wrong in
practice: the target is linked with --gc-sections and a version script, so
neither library ends up in the shared object's NEEDED list, and
conda-build reports both as overdepending on linux-64 and linux-aarch64:

    WARNING (onnxruntime-ep-tensorrt): run-exports library package
      libonnx==1.23.0 in requirements/run but it is not used
    WARNING (onnxruntime-ep-tensorrt): run-exports library package
      libabseil==20260526.0 in requirements/run but it is not used

libprotobuf and cuda-cudart-dev are not reported, so those two are
genuinely linked and stay. Comment added so they are not added back
without checking that report.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
<details><summary>Claude's draft</summary>

The runtime test had grown to 13 named checks over 493 lines -- import order
in subprocesses, six models at two optimization levels, contrib fusion, ORT
format, LoRA adapters, CoreML with profiling assertions, CUDA library
loading, CUDA on GPU. That is more than this branch needs to defend and
more than a reviewer can be expected to read.

All of it existed to catch one failure mode. Unvendoring onnx, protobuf and
abseil means onnx and onnxruntime now share one copy of the onnx protobuf
descriptors and one ONNX schema registry, and that is what breaks in ways
an import check cannot see: duplicate descriptor registration, duplicate or
missing schema registration, symbols libonnx does not export, exception
typeinfo that no longer matches across libraries. All of them show up in a
single scenario -- build a model with onnx, have onnxruntime run it, then
use onnx again -- which is the sequence that produced
conda-forge/onnx-feedstock#157, conda-forge#158 and conda-forge#159.

So there is now one flat script per codepath, 69 and 103 lines, no helper
layer:

  * test_onnxruntime_runtime.py imports onnxruntime before onnx (the order
    that broke), builds and checks and shape-infers with onnx, runs at
    ORT_ENABLE_ALL so the optimizers register contrib schemas into the
    shared registry, compares against onnx's reference evaluator on every
    provider the hardware allows, then re-checks the model with onnx.
  * cpp_interop/ does the same through the onnx C++ API and a second
    Ort::Env.

test_ep_tensorrt.py is 51 lines: a bad install of that package fails in
exactly two ways, wrong directory or unresolved libnvinfer, and loading the
library with RTLD_NOW from onnxruntime's search directory catches both
without a GPU, which is what CI can run. With a GPU the model run proves it
end to end.

Dropped rather than forgotten: ORT format and LoRA round trips, because the
flatbuffers schema headers carry a static_assert pinning the flatbuffers
version, so a bad regeneration fails the build instead; the subprocess
import-order matrix, because the surviving order is the only one that ever
broke; and CoreML profiling assertions, though CoreML still runs the
comparison with a looser bound since it may evaluate in fp16.

Verified against locally built CUDA 13.4 packages on an RTX 4090: the
python test passes on CPU and CUDA with ONNXRUNTIME_TEST_REQUIRE_CUDA=1,
the TensorRT test passes with ONNXRUNTIME_TEST_REQUIRE_TENSORRT=1 and fails
loudly in an environment without the package, and the C++ test was compiled
against the built onnxruntime-cpp and run.

Resume this Claude session:
```
cd /home/mark/git/feedstock/onnxruntime-feedstock
claude --resume 3cf41599-db98-4c6f-a67a-ee7dcacda8b4
```
</details>

Claude-Session: https://claude.ai/code/session_01Jxvsn2W6MfjGKXcabWVHNe
@hmaarrfk hmaarrfk changed the title Unvendor bundled C++ deps (protobuf, abseil, onnx, re2, flatbuffers, coremltools) Unvendor the bundled C++ deps (protobuf, abseil, onnx, re2, flatbuffers, coremltools), and add an optional TensorRT EP Sep 26, 2026
@hmaarrfk
hmaarrfk marked this pull request as ready for review September 27, 2026 18:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants