Skip to content

fix(qualification): accept semantic text task IDs in encoder parity - #1607

Closed
jkzhang7 wants to merge 1 commit into
NVIDIA:mainfrom
jkzhang7:fix/encoder-accuracy-semantic-tasks
Closed

jkzhang7 wants to merge 1 commit into
NVIDIA:mainfrom
jkzhang7:fix/encoder-accuracy-semantic-tasks

Conversation

@jkzhang7

@jkzhang7 jkzhang7 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Background

The shared STS accuracy qualification (embedding_vector_parity, benchmark
stsbenchmark_embedding_parity) chooses its Hugging Face reference mode and
benchmark operation from the candidate's task name, through a table in
_encoder_embedding_parity that only knows the retired encoding and
embedding names. A family migrated to the semantic Task SDK declares
text_to_pooled_features or text_to_embedding instead, so its qualification
fails before inference with:

QualificationError: embedding vector parity does not support candidate task 'text_to_pooled_features'

This was reported on #1584 (electra, mpnet and xlnet) by a reviewer, who
suggested a separate linked prerequisite so family migrations stay
self-contained. Every encoder family migrated to the semantic Tasks (including
#1584 and the other open encoder migrations) needs this routing.

Exit Criteria

  • _encoder_embedding_parity accepts text_to_pooled_features and
    text_to_embedding and routes them exactly like the retired names: the
    first-token reference mode with encode, and the mean-pool reference mode
    with embed.
  • The reference runner, gates and acceptance thresholds are unchanged.
  • The retired names still work for families that have not migrated.
  • Non-goal: any change to families, the benchmark worker, or the performance
    path (task_adapters.py already maps these Tasks).

Implementation

Affected component: shared qualification tooling only
(qualification_tests/benchmark_qualification/accuracy.py) and its test. No
family or runtime file is touched.

  • accuracy.py: two entries added to the task table:
    text_to_pooled_features -> ("cls", "encode") and
    text_to_embedding -> ("embedding", "embed").
  • tools/tests/test_model_benchmark.py: the existing parametrized
    test_encoder_accuracy_uses_task_semantics_without_model_specific_runner
    now also covers both semantic IDs, asserting the HF reference mode, the
    benchmark operation and the candidate requests.

The candidate side needs no change: for these Tasks the benchmark worker
already reports values, dim and feature_kind, which is what the
comparison reads.

Change categories

  • CI or developer tooling

Validation

Commands and Results

$ PYTHONPATH=core/builder:apps/benchmark:. python3 -m pytest tools/tests/test_model_benchmark.py -q -k "encoder_accuracy or hf_encoder"
6 passed

$ PYTHONPATH=core/builder:apps/benchmark:. python3 -m pytest tools/tests/test_architecture.py -q
55 passed

Without the accuracy.py change, the two new test cases fail with the error
quoted above; with it they pass. Two unrelated tests in
test_model_benchmark.py (test_hf_translation_supports_replace_final_unknown_tokenizer)
fail on this machine on the unmodified tree because torch is not installed.

Real hardware: the three STS accuracy qualification cases that failed on
#1584, run through the repository's own runner on an RTX 4090 (SM 8.9, driver
595.91, image nvcr.io/nvidia/tensorrt:26.07-py3, TensorRT 11.1.0.106), with
this change merged onto the #1584 branch (426509c9):

python -m tools.model_benchmark run --model <model> --kind accuracy \
  --dataset stsbenchmark-test=<STSBenchmark/stsbenchmark_test.jsonl> \
  --runtime-root <family runtime> --trtmc-bench <trtmc-bench> --worker <trtmc_benchmark_worker>

The dataset file was fetched from mteb/stsbenchmark-sts and its SHA-256
(7c8927ee...e5bb) matches the digest pinned in
stsbenchmark_embedding_parity.yaml. All gates are the benchmark's own,
unchanged (100 samples, 50 pairs):

model                       status  vector_pass_rate  min_vector_cosine  max_pair_cosine_abs_delta
electra-base-discriminator  passed  1.0               0.999974           0.001622
all-mpnet-base-v2           passed  1.0               0.999981           0.002321
xlnet-base                  passed  1.0               0.999990           0.000486

HF vs candidate STS Spearman: electra -0.3244 / -0.3238, mpnet 0.9408 /
0.9404, xlnet -0.1678 / -0.1678 (the electra and xlnet base checkpoints are not
sentence-embedding models, so their absolute Spearman is low; the gate is
parity with the Hugging Face reference, not quality).

On the pod, pip install -e . failed, so trtmc-bench was invoked through a
two-line wrapper around python -m trtmc_benchmark.cli and the built trtmc
was linked at core/builder/tensorrt_model_connect/bin/trtmc; neither affects
the code under test.

Hardware, Environment, and Revisions

  • Repository head: based on github/main at 5c675fb1.

Not Run / Remaining Gaps

  • The performance qualification path is not changed and was not run here.

Contributor Self-Review

  • I have completed a self-review of this change.

Notes For Future Readers

Risk level

  • Low

Two table entries in shared qualification tooling, covered by a test that fails
without them; the retired names are unchanged.

…g parity

The STS embedding_vector_parity accuracy path routed the candidate task
through a table that only knew the retired encoding and embedding names, so
a family migrated to text_to_pooled_features or text_to_embedding failed
before inference with "embedding vector parity does not support candidate
task". Map the semantic IDs to the same HF reference mode and benchmark
operation as the retired names (first-token reference with encode, mean-pool
reference with embed); reference and gates are unchanged, and the retired
names stay for families that have not migrated yet.

Extend the existing task-semantics test with both semantic IDs. The candidate
side needs no change: the benchmark worker already reports values, dim and
feature_kind for these Tasks.

Signed-off-by: Jingkun Zhang <jkzhang7@hotmail.com>
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@jkzhang7

jkzhang7 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Folded into #1584 (commit d7ddd74) so that PR is correct on its own; closing this one.

@jkzhang7 jkzhang7 closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant