Repository navigation
Conversation
…g parity The STS embedding_vector_parity accuracy path routed the candidate task through a table that only knew the retired encoding and embedding names, so a family migrated to text_to_pooled_features or text_to_embedding failed before inference with "embedding vector parity does not support candidate task". Map the semantic IDs to the same HF reference mode and benchmark operation as the retired names (first-token reference with encode, mean-pool reference with embed); reference and gates are unchanged, and the retired names stay for families that have not migrated yet. Extend the existing task-semantics test with both semantic IDs. The candidate side needs no change: the benchmark worker already reports values, dim and feature_kind for these Tasks. Signed-off-by: Jingkun Zhang <jkzhang7@hotmail.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
6 tasks done
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
The shared STS accuracy qualification (
embedding_vector_parity, benchmarkstsbenchmark_embedding_parity) chooses its Hugging Face reference mode andbenchmark operation from the candidate's task name, through a table in
_encoder_embedding_paritythat only knows the retiredencodingandembeddingnames. A family migrated to the semantic Task SDK declarestext_to_pooled_featuresortext_to_embeddinginstead, so its qualificationfails before inference with:
This was reported on #1584 (electra, mpnet and xlnet) by a reviewer, who
suggested a separate linked prerequisite so family migrations stay
self-contained. Every encoder family migrated to the semantic Tasks (including
#1584 and the other open encoder migrations) needs this routing.
Exit Criteria
_encoder_embedding_parityacceptstext_to_pooled_featuresandtext_to_embeddingand routes them exactly like the retired names: thefirst-token reference mode with
encode, and the mean-pool reference modewith
embed.path (
task_adapters.pyalready maps these Tasks).Implementation
Affected component: shared qualification tooling only
(
qualification_tests/benchmark_qualification/accuracy.py) and its test. Nofamily or runtime file is touched.
accuracy.py: two entries added to the task table:text_to_pooled_features -> ("cls", "encode")andtext_to_embedding -> ("embedding", "embed").tools/tests/test_model_benchmark.py: the existing parametrizedtest_encoder_accuracy_uses_task_semantics_without_model_specific_runnernow also covers both semantic IDs, asserting the HF reference mode, the
benchmark operation and the candidate requests.
The candidate side needs no change: for these Tasks the benchmark worker
already reports
values,dimandfeature_kind, which is what thecomparison reads.
Change categories
Validation
Commands and Results
Without the
accuracy.pychange, the two new test cases fail with the errorquoted above; with it they pass. Two unrelated tests in
test_model_benchmark.py(test_hf_translation_supports_replace_final_unknown_tokenizer)fail on this machine on the unmodified tree because
torchis not installed.Real hardware: the three STS accuracy qualification cases that failed on
#1584, run through the repository's own runner on an RTX 4090 (SM 8.9, driver
595.91, image
nvcr.io/nvidia/tensorrt:26.07-py3, TensorRT 11.1.0.106), withthis change merged onto the #1584 branch (
426509c9):The dataset file was fetched from
mteb/stsbenchmark-stsand its SHA-256(
7c8927ee...e5bb) matches the digest pinned instsbenchmark_embedding_parity.yaml. All gates are the benchmark's own,unchanged (100 samples, 50 pairs):
HF vs candidate STS Spearman: electra -0.3244 / -0.3238, mpnet 0.9408 /
0.9404, xlnet -0.1678 / -0.1678 (the electra and xlnet base checkpoints are not
sentence-embedding models, so their absolute Spearman is low; the gate is
parity with the Hugging Face reference, not quality).
On the pod,
pip install -e .failed, sotrtmc-benchwas invoked through atwo-line wrapper around
python -m trtmc_benchmark.cliand the builttrtmcwas linked at
core/builder/tensorrt_model_connect/bin/trtmc; neither affectsthe code under test.
Hardware, Environment, and Revisions
github/mainat5c675fb1.Not Run / Remaining Gaps
Contributor Self-Review
Notes For Future Readers
accuracy.py(two table entries), then the test parameters.Risk level
Two table entries in shared qualification tooling, covered by a test that fails
without them; the retired names are unchanged.