feat(eval): add translation evaluation and refresh opus-mt-en-ru recipes - #1394
ssss141414 wants to merge 3 commits into
Conversation
|
APPROVE No blocking findings. Reviewed exact candidate All 8 authoritative manifests and 378 members rehashed exactly. Preserved L0/L1 reuse is valid after a 38-file artifact rehash and no-impact proof. Four artifacts have the claimed signatures, opset 17, realized FLOAT/FLOAT16 dtypes, colocated external data, and memory-bearing perf: FP32 encoder/decoder Independent L2 rerun reproduced FP32 encoder/decoder cosine Reviewer reran focused Exact-head GitHub checks:
PR #1213 remains OPEN draft and conflicting at Disposition: terminal skill-level APPROVE. Keep the PR draft. This is an ordinary conversation comment, not a GitHub Review state. |
Summary
This adds generalized translation evaluation for WinML encoder-decoder models, refreshes the four existing Helsinki-NLP/opus-mt-en-ru CPU recipes, and fixes encoder cross-attention mask alignment during generation. The evaluator uses TorchMetrics corpus SacreBLEU with 13a tokenization and standard chrF2; no dependency is added. The contribution is Effort L2 / Outcome L2 and reaches L3 PASS with full planned CPU fp32/fp16 coverage; the two-row Eval is functional-smoke evidence only, not representative translation quality.
Model metadata
What the model does
Helsinki-NLP/opus-mt-en-ru accepts English text and autoregressively produces Russian text using a Marian encoder-decoder translation model. Evidence: pinned checkpoint
Helsinki-NLP/opus-mt-en-ruat revisionbb09c99d180016eac6819df3dae68edb1690fdee, whose Hub metadata identifies translation, English, Russian, and Apache-2.0; pinned config identifiesMarianMTModel,model_type=marian, andis_encoder_decoder=true. Confidence: verified.Primary user stories
Supported tasks
AutoModelForSeq2SeqLMmapping, and WinML composite registryMarianOnnxConfigvendor registry and WinMLMarianDecoderIOConfigoverrideMarianOnnxConfigvendor registry and WinMLMarianEncoderIOConfigoverrideModel architecture
MarianMTModelsource, and fresh HTP/ONNX artifacts (verified).Validation and support evidence
1. Baseline
The baseline was fully rerun on
maincommitbe3e59dd412d4918c5a852aa4d8f0207c34aaf6fwith WinML0.0.1.dev0; broad runtime, dependency, Marian, generation, Eval, and Analyze changes since the prior baseline made reuse unsafe. CPU fp32 auto-build passed in 137.19 s and emitted opset-17 encoder and cached-decoder artifacts at fixed batch 1, source/cache length 512, and vocabulary 62,518. Baseline perf passed: encoder average/p50 69.76/69.93 ms, 14.34 samples/s, RAM +64.4 MB; decoder average/p50 21.68/20.78 ms, 46.12 samples/s, RAM +21.2 MB. Encoder parity passed on a pinned real row with cosine0.9999999999993413and max absolute error5.245208740234375e-06.Baseline two-row generation mechanically ran but collapsed both inputs to the same low-quality short Russian output. Translation Eval was
UNSUPPORTED-TASKbefore dataset loading; an ad hoc probe returned SacreBLEU-13a0.0and chrF20.012449raw (1.2449366on the 0-100 scale), as functional smoke only. Analyze produced parseable partial-success JSON: encoder 197 nodes / 15 operator types and decoder 378 nodes / 20 operator types. Optimum exposed feature-extraction/text-generation/text2text-generation with and without past; WinML added no keys but replaced vendor feature-extraction and text2text-generation configurations with Marian-specific I/O overrides (VENDOR+OVERRIDE).2. Goal
No downstream role silently re-tiered the frozen charter.
3. Outcome
L3 PASS, full planned coverage, no deferred tuples. The final candidate is
3f5f321e4721e3022f7f50514e8cfcfb9ad4570d(treeddf4c7810f2029937798e3ca1c03e4e6595a07af) on basebe3e59dd412d4918c5a852aa4d8f0207c34aaf6f; it is clean and unpushed in the sealed handoff.Shipped product paths comprise the four existing Model25 CPU recipes plus generalized Eval/runtime code and focused tests. There is no hardcoded checkpoint-name branch, no dependency addition, no README change, and no duplicate recipe path. Learner findings
marian-010throughmarian-013and methodology findings_meta-117and_meta-118were published separately in Lane A draft PR #314 at commit2f4c04fa1ff3b427d8bb9c39f7c95c323d178ec3; no skill files are included in this model contribution.Quality gates
4. Per-EP/device/precision results and Functional smoke Eval
Goal ladder and artifact signatures
L0 and L1 are PRESERVED, not re-executed on the final runtime-repair SHA. They executed on
92387df1fbe50b628ac38574cdcb0fa992247088and were reused only after 38 sealed artifact files and all recipes rehashed exactly and the repair was proven evaluator/runtime-only with no config, export, optimizer, recipe, lock, artifact-input, or component-perf impact. The execution SHA remains92387df1...; it is not relabeled as final-candidate execution.input_ids[1,512],attention_mask[1,512]; outputencoder_hidden_states; model 67,274 bytes +model.onnx.data204,742,656 bytes; initializers FLOAT x102, INT64 x29; opset 17decoder_input_ids[1,1],encoder_hidden_states[1,512,512],attention_mask[1,512],decoder_attention_mask[1,512],cache_position[1], 12 KV tensorspast_{0..5}_{key,value}[1,8,512,64]; outputslogits+ 12present_{0..5}_{key,value}; model 146,157 bytes +model.onnx.data358,291,672 bytes; initializers FLOAT x166, INT64 x43; opset 17input_ids[1,512],attention_mask[1,512]; outputencoder_hidden_states; model 102,432,639 bytes; initializers FLOAT16 x102, INT64 x29; opset 17model.onnx.data179,156,076 bytes; initializers FLOAT16 x166, INT64 x43; opset 17All four artifact checks also passed co-located external-data, vocabulary 62,518, and realized-precision validation.
Perf
L2 parity and generation
0.9999999999993413/5.245208740234375e-060.9999999999986515/1.71661376953125e-050.999998659921451/0.0079450607299804690.9999992557247197/0.014481067657470703For both fp32 and fp16, ONNX and pinned PyTorch produced these exact, distinct full greedy sequences and decoded texts:
[62517, 71, 2071, 575, 497, 221, 11, 20840, 30, 164, 3427, 2, 119, 26, 3712, 41, 2207, 806, 8432, 1300, 2, 119, 3983, 234, 38192, 53, 4410, 1046, 17820, 146, 3, 0]Text:
"У нас есть 4-месячные мыши, которые не диабиотичны, которые раньше были диабетиками", добавил он.[62517, 1094, 5192, 1432, 21824, 261, 779, 2, 13172, 25978, 443, 374, 21647, 3085, 48, 1974, 10908, 6, 37750, 43669, 30, 2, 10383, 1561, 806, 6959, 847, 2, 7, 15382, 42, 4421, 10997, 7, 22526, 6876, 40306, 48, 7215, 38192, 41, 28840, 53, 2, 28, 2549, 145, 401, 4275, 6, 4790, 1029, 11078, 3, 0]Text:
Доктор Эхуд Ур, профессор медицины Далхаузийского университета в Галифаксе, Новая Шотландия, и председатель клинического и научного отделения Канадской ассоциации диабета предупредили, что исследования все еще находятся в начале своего существования.Functional smoke Eval
PASS on final candidate
3f5f321e4721e3022f7f50514e8cfcfb9ad4570d, FP32 CPU only. Before this contribution, WinML rejectedtranslationas an unsupported Eval task. The generalized evaluator now performs explicit source/reference extraction, bounded composite generation, decoding, exact accounting, and corpus metrics through TorchMetricsSacreBLEUScore(tokenize='13a')andCHRFScore(n_char_order=6, n_word_order=0, beta=2.0)on a 0-100 scale.gsarti/flores_101source revisionbc58ae43b22607b3e1e2bf3ae1bc5cb053495abb, immutable Parquet revision1a45e707ea4b5ea3d6c71341f18bfca6a6e356a3, configall, splitdevtest, deterministic rows 1-2,sentence_eng->sentence_rus.9.7950; chrF256.8509. Independent recomputation returned the same values."We now have 4-month-old mice that are non-diabetic that used to be diabetic," he added."Теперь у нас есть четырёхмесячные мыши, у которых больше нет диабета", — добавил он."У нас есть 4-месячные мыши, которые не диабиотичны, которые раньше были диабетиками", добавил он.Dr. Ehud Ur, professor of medicine at Dalhousie University in Halifax, Nova Scotia and chair of the clinical and scientific division of the Canadian Diabetes Association cautioned that the research is still in its early days.Согласно предупреждению доктора Эхуда Ура (Ehud Ur), профессора медицины в Университете Делхаузи в Галифаксе (Новая Шотландия) и председателя клинико-научного отдела Канадской диабетической ассоциации, исследования все еще находятся на начальной стадии.Доктор Эхуд Ур, профессор медицины Далхаузийского университета в Галифаксе, Новая Шотландия, и председатель клинического и научного отделения Канадской ассоциации диабета предупредили, что исследования все еще находятся в начале своего существования.This proves end-to-end evaluator operability only. Two adjacent rows from one document are not representative sampling, and these values are not benchmark-quality or representative accuracy claims. Eval was not repeated for fp16 and no fp16 accuracy claim is made.
5. Delta
Commit delta
caa7dfb972b0c8cc1c5c878defcde790139eaf3292387df1fbe50b628ac38574cdcb0fa9922470883f5f321e4721e3022f7f50514e8cfcfb9ad4570dattention_maskequivalentlyChanged files
examples/recipes/Helsinki-NLP_opus-mt-en-ru/cpu/cpu/translation_fp16_decoder_config.jsonexamples/recipes/Helsinki-NLP_opus-mt-en-ru/cpu/cpu/translation_fp16_encoder_config.jsonexamples/recipes/Helsinki-NLP_opus-mt-en-ru/cpu/cpu/translation_fp32_decoder_config.jsonexamples/recipes/Helsinki-NLP_opus-mt-en-ru/cpu/cpu/translation_fp32_encoder_config.jsonsrc/winml/modelkit/eval/__init__.pysrc/winml/modelkit/eval/evaluate.pysrc/winml/modelkit/eval/metrics/translation.pysrc/winml/modelkit/eval/translation_evaluator.pysrc/winml/modelkit/models/winml/encoder_decoder.pysrc/winml/modelkit/utils/eval_utils.pytests/unit/eval/test_translation_evaluator.pytests/unit/models/winml/test_composite_from_pretrained.pyAll four existing recipes change only
/export/compatibility/transformers_attentionfrom absent (null) to"eager", matching current auto-config; fp16 quantization is preserved. The evaluator registry is additive; translation direction comes from explicit source/reference fields, bounded capacities come from model metadata, metrics are corpus-aggregated using the existing TorchMetrics dependency, failures propagate, and zero evaluated rows fails closed.examples/recipes/README.mdremains untouched.Bug fix explanation
_run_decoderapplied cache-oriented left-padding to every decoder feed. This moved each short encoder attention mask from active source positions0..30/0..47to481..511/464..511, so decoder cross-attention ignored the actual source tokens. Already-512-wide one-step parity inputs hid the defect.WinMLEncoderDecoderModelnow preserves right-padding for the encoder cross-attention mask while retaining left-padding where the static decoder cache requires it; the regression exercisestest_encoder_decoder_preserves_encoder_alignment_across_independent_rows.0..30/0..47; two source rows produce distinct encoder states and distinct outputs; full bounded fp32 and fp16 ONNX token sequences and decoded translations exactly match pinned PyTorch.PR #1213 overlap
As last checked, microsoft/winml-cli PR #1213 was OPEN, DRAFT, DIRTY/CONFLICTING at head
a74eb377a2d43468c047ea10649aa52faec48991; no Model25 PR existed. It overlaps the generalized evaluator. If #1213 merges first, dropcaa7dfb9, rebase the recipe/runtime commits onto currentmain, return to Planner for dependency-aware impact analysis, and rerun Eval plus every invalidated check. Do not drop3f5f321eunless the merged encoder-decoder runtime proves equivalent encoder-mask alignment.6. Analyze summary — component level and op level
Static rule analysis is ANALYZE-PARTIAL-SUCCESS: both all-EP commands emitted complete seven-target JSON, while exit 1 reflects providers without shipped rule data. This is compatibility analysis, not runtime execution.
Component-level summary
Op-level summary
Provider absence and missing rule data are distinct from an unsupported verdict. No accelerator runtime support is claimed from static rules.
7. Reproduce commands