feat: register nemotron-rerank-vl-1b-v2 alongside the existing reranker - #2553
Draft
erichare wants to merge 3 commits into
Draft
feat: register nemotron-rerank-vl-1b-v2 alongside the existing reranker#2553erichare wants to merge 3 commits into
erichare wants to merge 3 commits into
Conversation
Adds nvidia/llama-nemotron-rerank-vl-1b-v2 (NIM 2.x, Rust runtime) as a second, non-default model. The existing llama-3.2-nv-rerankqa-1b-v2 stays is-default: collections pin provider + model in their schema, so the default only affects newly created collections and existing traffic is unaffected. Batch size is held at 10 to match the existing model so load comparisons are like-for-like; NIM 2.x accepts up to 512 passages per call, which is worth measuring as a separate change. Note this file is only used when the embedding gateway is disabled (RerankingProviderConfigProducer) - Astra dev/prod get the model list from EGW over gRPC, so prod registration is an EGW config change. This entry covers local/OSS/HCD and local testing against the new NIM.
Contributor
➡️ Unit Test Coverage Delta vs Main Branch
|
Contributor
Unit Test Coverage Report
|
Contributor
➡️ Integration Test Coverage Delta vs Main Branch (dse69-it)
|
Contributor
Integration Test Coverage Report (dse69-it)
|
Contributor
➡️ Integration Test Coverage Delta vs Main Branch (hcd-it)
|
Contributor
Integration Test Coverage Report (hcd-it)
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
nvidia/llama-nemotron-rerank-vl-1b-v2(NIM 2.x, the new Rust runtime) as a second, non-default reranking model, so it can run alongsidenvidia/llama-3.2-nv-rerankqa-1b-v2rather than replacing it.Test plan
test-reranking-providers-config.yaml)findRerankingProviderslists both models, old one flagged defaultfindAndRerankagainst a NIM 2.x pod using the new model