Skip to content

Pull requests: huggingface/text-embeddings-inference

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Cuda flake fix
#913 opened Aug 10, 2026 by banderlog Loading…
Fix build via nix-shell
#912 opened Aug 10, 2026 by banderlog Loading…
Add OTLP over HTTP/protobuf export support
#910 opened Aug 4, 2026 by AmreshSinha Loading…
5 tasks
docs: note OpenAI client base_url for multi-model gateways
#905 opened Aug 3, 2026 by seven7763 Loading…
3 tasks
Add SigLIP model - text embeddings only
#903 opened Jul 28, 2026 by m-toman Draft
4 of 5 tasks
Speed up ModernBERT inference by 1.15x on CUDA
#897 opened Jul 11, 2026 by hotchpotch Loading…
2 of 4 tasks
v1.9.4
Add --root-path to specificy HTTP prefix
#894 opened Jul 9, 2026 by alvarobartt Member Draft
4 of 5 tasks
v1.9.4
feat(qwen3): support reranker on candle backend
#886 opened Jun 26, 2026 by malaiwah Loading…
4 of 5 tasks
feat: allow controlling startup warmup tokens
#884 opened Jun 25, 2026 by malaiwah Loading…
4 of 5 tasks
Add HTTP rerank endpoint coverage
#878 opened Jun 22, 2026 by jesco-absolut Loading…
support b300
#875 opened Jun 13, 2026 by deepindeed2022 Loading…
Upgrade hf-hub to 1.0.0-rc.1
#865 opened May 7, 2026 by assafvayner Draft
2 of 4 tasks
feat: better prometheus buckets for batch_size, isl, and input length
#847 opened Mar 20, 2026 by michaelfeil Contributor Loading…
5 tasks
v1.10.0
Support Qwen3-Reranker
#835 opened Feb 20, 2026 by kozistr Contributor Loading…
4 of 5 tasks
v1.10.0
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.