You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This adds CPU fp32 and fp16 reranking recipes for Alibaba-NLP/gte-reranker-modernbert-base, an English cross-encoder that emits one raw relevance logit per query-document pair. The shipped Effort and Outcome are L2, with model-specific recipe configuration layered on the generic reranking capability owned by dependency PR #1322. Sealed candidate evidence reaches L3 PASS with full coverage, and exact candidate-SHA quality evidence passes all required GitHub Actions gates.
Model metadata
What the model does
An English cross-encoder text reranker that jointly encodes a query and candidate document and emits one raw relevance logit per pair for descending-score ranking.
Evidence: the pinned model card identifies the model as a text reranker and demonstrates pair scoring; its pinned config.json declares model_type=modernbert, ModernBertForSequenceClassification, num_labels=1, and classifier_pooling=mean.
Confidence: verified.
Primary user stories
A user supplies a query and candidate documents to obtain one relevance score per query-document pair and reorder retrieval results for search or RAG. Evidence: pinned model card pair-scoring example. Confidence: verified.
A user supplies an English question and passages to obtain a higher score for the more relevant passage in a bounded reranking stage. Evidence: pinned model card pair-scoring example. Confidence: verified.
Supported tasks
reranking on the checkpoint surface. Evidence: pinned model card model type and pair-scoring example. Confidence: verified.
text-classification on the Transformers, Optimum ONNX, and WinML surfaces. Evidence: pinned config.json; Optimum's ModernBERT vendor tasks include text classification; WinML recipe-free export resolved text classification successfully. Confidence: verified.
Source/confidence: pinned checkpoint config, selected Transformers ModernBertForSequenceClassification source, and exported hierarchy metadata (verified). The encoder-only cross-encoder has 149,605,633 parameters; the export traced 22 encoder layers and produces scalar logits.
Validation and support evidence
1. Baseline
Pinned main: 0876e5ae1c98a169a6137e092e0d7b30bf9cee33; WinML version 0.0.1.dev0.
Recipe-free auto-config resolved text-classification. Optimum advertised feature-extraction, fill-mask, text-classification, and token-classification; WinML added no vendor export task. Reranking is a WinML semantic over the same single-logit sequence-classification export supplied by dependency PR Add WinML reranking support for cross-encoder/ms-marco-MiniLM-L6-v2 #1322.
Build: PASS, ModernBertForSequenceClassification, task text-classification, opset 17, logits output, completed in 81.2 s; an independent repeat also passed.
Perf floor: 941.98 ms mean, 941.60 ms p50, 975.66 ms p95, 1.06 samples/s, and +380.8 MB RAM on Intel Core i9-10900X; input [1,1024], output logits [1,1] float32.
Eval floor: exit 1 because main resolved the requested task string but reported that reranking was not supported by WinML eval.
Goal floor: L1. Baseline evidence was freshly rerun at the pinned commit; no prior stages were reused.
2. Goal
Committed Effort: L2.
Committed Goal ceiling: L3.
Committed Outcome: L2.
Success definition: build the required CPU/cpu fp32 and fp16 tuples; report CPU perf; establish PyTorch-to-ONNX raw-logit parity and descending-order agreement on pinned model-card pairs; and run one bounded fp32 CPU SciDocs reranking functional smoke with explicit group, candidate, pair, and sequence caps.
No ceiling change or re-issued charter occurred. The sealed Goal result reached L3 PASS with full coverage and no deferred tuples.
3. Outcome
The model-specific Outcome is L2: exactly two checked-in CPU reranking recipes select dependency-owned generic reranking behavior without adding shared code. The highest reached Goal verdict is L3 PASS, sealed for candidate 5833e0f8bb3d2b3df95cf813861c7692fa6459af, with full coverage and no deferred tuples.
Learner findings modernbert-001 through modernbert-004 retain the fp32/fp16 export, perf, parity, SciDocs smoke, and analyze evidence. Methodology finding _meta-112 records that quality conclusions require a candidate-local interpreter, tool, import root, and exact-lock provenance; its Lane A changes are separate from this model PR (compare).
Current candidate quality status: reentry8 is PASS for exact candidate 5833e0f8bb3d2b3df95cf813861c7692fa6459af. The authoritative exact-head validation run passed all six jobs: analyze 1,526 passed / 45 skipped; optim 848 passed / 16 skipped / 1 xfailed; commands 3,641 passed / 9 skipped / 1 warning; models 1,534 passed / 6 skipped / 2 xfailed; remaining 871 passed / 2 skipped / 1 deselected / 1 warning; and lint license headers and Ruff PASS with mypy clean across 438 source files. Every one of the six jobs explicitly checked out ssss141414/winml-cli at immutable ref 5833e0f8bb3d2b3df95cf813861c7692fa6459af, asserted HEAD, used the candidate-local .venv, and synced against the locked uv.lock SHA-256 F6672F52199FA442A7B1414BA943C4ECCC365940206E430777A51A1374B29053. The draft validation workflow PR #1337 commit is not the model candidate and must never merge. Goal-evidence integrity was independently repaired with an external immutable full SHA-256 snapshot seal covering all referenced Goal roots (226 source files total, zero mismatches); the sealed Goal L0-L3 and analysis values remain unchanged.
4. Per-EP/device/precision results and Functional smoke Eval
Tier
EP / Device
Precision
Verdict
Mean
p50
p95
Throughput
RAM delta
L0
CPUExecutionProvider / cpu
fp32
PASS
-
-
-
-
-
L0
CPUExecutionProvider / cpu
fp16
PASS
-
-
-
-
-
L1
CPUExecutionProvider / cpu
fp32
PASS
943.47 ms
943.94 ms
981.87 ms
1.06 samples/s
+389.2 MB
L1
CPUExecutionProvider / cpu
fp16
PASS
1294.34 ms
1290.29 ms
1338.52 ms
0.77 samples/s
+406.4 MB
L0 fp32: input_ids and attention_mask INT32 [1,1024]; FLOAT logits [1,1]; 193 FLOAT, 15 INT64, and 2 BOOL initializers; external data 601,194,496 bytes.
L0 fp16: the same input/output contract; 193 FLOAT16, 15 INT64, and 2 BOOL initializers; external data 301,126,144 bytes. The fp16/fp32 external-data ratio is 0.5008797419196599.
L2 fp32, three pinned model-card pairs: cosine 0.9999999999996143, max absolute difference 2.6226043701171875e-06; PyTorch and ONNX descending order both [1,0,2], so order matched. Scores were untransformed raw scalar logits.
L2 fp16, the same three pairs: cosine 0.9999993407707104, max absolute difference 0.004093289375305176; PyTorch and ONNX descending order both [1,0,2], so order matched. Scores were untransformed raw scalar logits.
Functional smoke Eval: L3 PASS on final candidate 5833e0f8bb3d2b3df95cf813861c7692fa6459af, FP32 CPU, using mteb/scidocs-reranking test data pinned to revision 56a6d0140cf6356659e2a7c1413286a774468d44. Deterministic first-N streaming selection processed 2 query groups and 20 query-passage pairs. Caps were 2 groups, 10 candidates per group, 20 total pairs, and sequence length 1024. The schema used query, positive, and negative; positives mapped to relevance 1, negatives to 0; exactly one raw float logit per pair was ranked descending without categorical mapping or softmax. MRR@10 = 1.0, Recall@1 = 1.0, and Recall@10 = 1.0. This proves end-to-end operability only; it is not representative accuracy or benchmark quality. FP16 eval was not run because the Tester contract calls for one representative final-SHA FP32 CPU functional smoke. The former blocker was WinML's lack of a reranking evaluator; dependency PR #1322 provides generic task normalization, single-logit evaluation, grouped dataset adaptation, and ranking metrics.
5. Delta
The candidate diff contains exactly these two model-specific recipe files:
Pinned bounded SciDocs test plan: revision 56a6d0140cf6356659e2a7c1413286a774468d44, streaming true, shuffle false, samples 2, max candidates 10
fp32
/export/compatibility
{"transformers_attention":"eager"}
absent in checked-in source; auto-config resolves to eager
fp16
/loader/task
text-classification
reranking
fp16
/quant
null
Existing-model fp16 quant block with mode=fp16 and task=reranking
fp16
/eval
absent
Pinned bounded SciDocs test plan: revision 56a6d0140cf6356659e2a7c1413286a774468d44, streaming true, shuffle false, samples 2, max candidates 10
fp16
/export/compatibility
{"transformers_attention":"eager"}
absent in checked-in source; auto-config resolves to eager
Both recipes preserve AutoModelForSequenceClassification, modernbert, opset 17, eager attention compatibility, INT32 [1,1024] inputs, and logits output. The delta is reducibility-consistent with the charter. Recipe-free acceptance passed: published pipeline_tag=text-ranking resolves generically to canonical reranking with no model-ID hard-coding. Generic source code is owned by dependency PR #1322 at parent 3708969b731425b0c6d4b97920d1b5e6519bb013; this candidate has no code changes. Existing text-classification recipes and examples/recipes/README.md remain untouched.
6. Analyze summary - component level and op level
Static rule analysis completed with ANALYZE-PARTIAL-SUCCESS and exit code 1; this is compatibility analysis, not runtime execution. The unresolved overview gap is per-operator input-signature grouping: architecture regions are mapped through hierarchy scopes and tensor boundaries, while an ops-depth breakdown would be needed to resolve that finer grouping.
The mapped regions cover HTP-tagged attention and MLP scopes across all 22 encoder layers, an 11-node prediction-head region, and the scalar classifier tensor boundary. Embeddings, rotary embeddings, and mean pooling remain architecture-verified overview regions rather than operator-signature mappings.
Op-level summary
Artifact
Graph
Dominant ops
EP roll-up
fp32
948 operators / 23 types
MatMul 133; Mul 133; Slice 132; Add 110; Transpose 110
QNN partial: Gather, ReduceSum; Cast/Slice/Where unknown on rule-backed EPs
NvTensorRTRTX, OpenVINO GPU, and QNN GPU had rule-backed classifications but no runtime support on the test host; QNN's actionable partial types were OP/ai.onnx/Gather and OP/ai.onnx/ReduceSum. CUDA, MIGraphX, TensorRT, and DML had no populated rule classification in this handoff. No unsupported operator type was reported.
7. Reproduce commands
Set $MODEL to the pinned checkpoint directory and $OUT to a writable output directory. These are the tester-recorded build, perf, eval, and analyze command semantics with portable paths.
python -m winml.modelkit.cli build -c examples/recipes/Alibaba-NLP_gte-reranker-modernbert-base/cpu/cpu/reranking_fp32_config.json -m $MODEL-o "$OUT/fp32"
python -m winml.modelkit.cli build -c examples/recipes/Alibaba-NLP_gte-reranker-modernbert-base/cpu/cpu/reranking_fp16_config.json -m $MODEL-o "$OUT/fp16"--precision fp16
python -m winml.modelkit.cli perf -m "$OUT/fp32/model.onnx"--ep cpu --device cpu
python -m winml.modelkit.cli perf -m "$OUT/fp16/model.onnx"--ep cpu --device cpu
python -m winml.modelkit.cli eval -m "$OUT/fp32/model.onnx"--model-id $MODEL--task reranking --dataset mteb/scidocs-reranking --dataset-revision 56a6d0140cf6356659e2a7c1413286a774468d44 --split test --streaming --no-shuffle --samples 2--column query_column=query --column positive_column=positive --column negative_column=negative --column max_candidates=10--ep cpu --device cpu -o "$OUT/scidocs-fp32.json"--overwrite
python -m winml.modelkit.cli analyze --model "$OUT/fp32/model.onnx"--ep all --output "$OUT/fp32-analyze.json"
Reviewed exact candidate 5833e0f8bb3d2b3df95cf813861c7692fa6459af against dependency base/parent 3708969b731425b0c6d4b97920d1b5e6519bb013. Keep this PR in DRAFT.
Blocking findings:
Tester + explainer (_meta-112): Actions run metadata for 32647273077/32647273097 names the candidate SHA, but every retained grading log actually checks out synthetic merge commit 0752a41 (Merge 5833e0f8... into 0876e5ae...) and creates .venv under that merge checkout. The executed CI command also includes tests/unit/serve, while the captured candidate workflow omits it. Therefore candidate-local-provenance.v1.json and the PR-body statement that each job checked out the exact candidate and used its workflow are false. Rerun the required lint/mypy/full matrix from an exact candidate checkout with candidate-local locked hydration and captured interpreter/tool/import-root/version provenance, or report exact-candidate quality as NOT-RUN / ENVIRONMENT-BLOCKED; then have the explainer replace the upgraded PASS prose with tester-owned facts. Green check conclusions alone do not repair this provenance mismatch.
Tester evidence integrity: independent rehash found zero mismatches for the available 48-file terminal manifest, 4-file reentry3-final manifest, and 24-file reentry7 manifest. However, reentry3-final seals only four derived files; its finalizer reads L0/L1/L2/analyze evidence from tester-5833e0f8-20260823T034500Z-reentry3, which has no hash manifest/seal or independent rehash. Produce a fresh immutable Goal handoff that seals and independently rehashes every referenced acceptance-evidence root, preserving the current root as an incident record.
Checks that passed: live PR is OPEN, MERGEABLE, DRAFT, labeled model-scale-by-skill; head/base and parent are exact; dependency PR #1322 still points to the exact parent; diff is exactly the two CPU recipe JSON files with no README/shared-source/test delta; recipes match the existing exact-model text-classification shapes and fp16 convention while selecting generic dependency-owned reranking with no model hardcoding; independent ONNX validation confirmed IR 8/opset 17, INT32 [1,1024] inputs, and FLOAT [1,1] logits; frozen L0-L3 values, fp16 size/initializer reality, raw-logit parity/order, bounded pinned SciDocs accounting, and partial-success analysis agree with the retained artifacts; Lane A _meta-112 edits are separately pushed and absent from this model PR.
Reviewed exact candidate 5833e0f8bb3d2b3df95cf813861c7692fa6459af against dependency base/parent 3708969b731425b0c6d4b97920d1b5e6519bb013. Keep this PR in DRAFT.
Threads: pagination-aware enumeration found 1 review thread / 2 comments / 0 open; no next thread or comment pages. CodeQL thread recipe(gte-reranker-modernbert-base): add CPU reranking recipes #1336 (comment) is isResolved=true. The rationale is technically valid: the TYPE_CHECKING import supplies the torch.Tensor cast annotation, while _score_pair separately imports torch lazily for isinstance and torch.no_grad().
Evidence integrity: the external Goal seal covers the complete old matrix root (170 files, 7,222,063,602 bytes), terminal L3 root (50 files, 55,379 bytes), and reentry3-final root (6 files, 32,127 bytes): 226 source files, 7,222,151,108 bytes, full SHA-256, zero mismatches, unchanged pre/post metadata. Cross-root validation is PASS with exact candidate/base identities and unchanged L0-L3 values. Reentry8 terminal state is semantically consistent (PASS/completed/exit 0/no timeout/no error), its own 17-file seal independently rehashes with zero mismatches, and it supersedes the earlier quality-only handoff without altering Goal evidence.
Whole-PR review: OPEN, MERGEABLE/CLEAN, DRAFT, model-scale-by-skill; exact two-file recipe-only diff and production README unchanged. Current main remains the frozen baseline 0876e5ae1c98a169a6137e092e0d7b30bf9cee33; dependency PR #1322 remains at the exact base. Body hierarchy, source ownership, architecture/value fidelity, portable commands, delta, Analyze partial-success details, and no-local-path requirement pass. Both CPU tuples pass L0/L1; fp16 is a real 193-FLOAT16 artifact at 0.5008797 of fp32 external-data size. L2 raw-logit parity/order passes for fp32/fp16. The bounded pinned SciDocs FP32 CPU smoke passes L3 with 2 groups/20 pairs, MRR@10 1.0, Recall@1 1.0, Recall@10 1.0, explicitly operability-only. Coverage is full with no deferred tuples. Lane A _meta-112 implementation/finalization commits 8393e9b... / 5b58f0d... are present on the pushed learner branch and remain outside this model PR.
Residual blockers: none. Leave the explainer-owned shipment in DRAFT.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
model-scale-by-skillModel support PR created or maintained by the adding-model-support skill
2 participants
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This adds CPU fp32 and fp16 reranking recipes for
Alibaba-NLP/gte-reranker-modernbert-base, an English cross-encoder that emits one raw relevance logit per query-document pair. The shipped Effort and Outcome are L2, with model-specific recipe configuration layered on the generic reranking capability owned by dependency PR #1322. Sealed candidate evidence reaches L3 PASS with full coverage, and exact candidate-SHA quality evidence passes all required GitHub Actions gates.Model metadata
What the model does
An English cross-encoder text reranker that jointly encodes a query and candidate document and emits one raw relevance logit per pair for descending-score ranking.
config.jsondeclaresmodel_type=modernbert,ModernBertForSequenceClassification,num_labels=1, andclassifier_pooling=mean.verified.Primary user stories
verified.verified.Supported tasks
rerankingon the checkpoint surface. Evidence: pinned model card model type and pair-scoring example. Confidence:verified.text-classificationon the Transformers, Optimum ONNX, and WinML surfaces. Evidence: pinnedconfig.json; Optimum's ModernBERT vendor tasks include text classification; WinML recipe-free export resolved text classification successfully. Confidence:verified.Model architecture
ModernBertForSequenceClassificationsource, and exported hierarchy metadata (verified). The encoder-only cross-encoder has 149,605,633 parameters; the export traced 22 encoder layers and produces scalar logits.Validation and support evidence
1. Baseline
main:0876e5ae1c98a169a6137e092e0d7b30bf9cee33; WinML version0.0.1.dev0.text-classification. Optimum advertisedfeature-extraction,fill-mask,text-classification, andtoken-classification; WinML added no vendor export task. Reranking is a WinML semantic over the same single-logit sequence-classification export supplied by dependency PR Add WinML reranking support for cross-encoder/ms-marco-MiniLM-L6-v2 #1322.ModernBertForSequenceClassification, tasktext-classification, opset 17, logits output, completed in 81.2 s; an independent repeat also passed.[1,1024], output logits[1,1]float32.mainresolved the requested task string but reported thatrerankingwas not supported by WinML eval.2. Goal
3. Outcome
The model-specific Outcome is L2: exactly two checked-in CPU reranking recipes select dependency-owned generic reranking behavior without adding shared code. The highest reached Goal verdict is L3 PASS, sealed for candidate
5833e0f8bb3d2b3df95cf813861c7692fa6459af, with full coverage and no deferred tuples.Learner findings
modernbert-001throughmodernbert-004retain the fp32/fp16 export, perf, parity, SciDocs smoke, and analyze evidence. Methodology finding_meta-112records that quality conclusions require a candidate-local interpreter, tool, import root, and exact-lock provenance; its Lane A changes are separate from this model PR (compare).Current candidate quality status: reentry8 is PASS for exact candidate
5833e0f8bb3d2b3df95cf813861c7692fa6459af. The authoritative exact-head validation run passed all six jobs: analyze 1,526 passed / 45 skipped; optim 848 passed / 16 skipped / 1 xfailed; commands 3,641 passed / 9 skipped / 1 warning; models 1,534 passed / 6 skipped / 2 xfailed; remaining 871 passed / 2 skipped / 1 deselected / 1 warning; and lint license headers and Ruff PASS with mypy clean across 438 source files. Every one of the six jobs explicitly checked outssss141414/winml-cliat immutable ref5833e0f8bb3d2b3df95cf813861c7692fa6459af, assertedHEAD, used the candidate-local.venv, and synced against the lockeduv.lockSHA-256F6672F52199FA442A7B1414BA943C4ECCC365940206E430777A51A1374B29053. The draft validation workflow PR #1337 commit is not the model candidate and must never merge. Goal-evidence integrity was independently repaired with an external immutable full SHA-256 snapshot seal covering all referenced Goal roots (226 source files total, zero mismatches); the sealed Goal L0-L3 and analysis values remain unchanged.4. Per-EP/device/precision results and Functional smoke Eval
input_idsandattention_maskINT32[1,1024]; FLOAT logits[1,1]; 193 FLOAT, 15 INT64, and 2 BOOL initializers; external data 601,194,496 bytes.0.5008797419196599.0.9999999999996143, max absolute difference2.6226043701171875e-06; PyTorch and ONNX descending order both[1,0,2], so order matched. Scores were untransformed raw scalar logits.0.9999993407707104, max absolute difference0.004093289375305176; PyTorch and ONNX descending order both[1,0,2], so order matched. Scores were untransformed raw scalar logits.Functional smoke Eval: L3 PASS on final candidate
5833e0f8bb3d2b3df95cf813861c7692fa6459af, FP32 CPU, usingmteb/scidocs-rerankingtest data pinned to revision56a6d0140cf6356659e2a7c1413286a774468d44. Deterministic first-N streaming selection processed 2 query groups and 20 query-passage pairs. Caps were 2 groups, 10 candidates per group, 20 total pairs, and sequence length 1024. The schema usedquery,positive, andnegative; positives mapped to relevance 1, negatives to 0; exactly one raw float logit per pair was ranked descending without categorical mapping or softmax. MRR@10 =1.0, Recall@1 =1.0, and Recall@10 =1.0. This proves end-to-end operability only; it is not representative accuracy or benchmark quality. FP16 eval was not run because the Tester contract calls for one representative final-SHA FP32 CPU functional smoke. The former blocker was WinML's lack of a reranking evaluator; dependency PR #1322 provides generic task normalization, single-logit evaluation, grouped dataset adaptation, and ranking metrics.5. Delta
The candidate diff contains exactly these two model-specific recipe files:
examples/recipes/Alibaba-NLP_gte-reranker-modernbert-base/cpu/cpu/reranking_fp32_config.jsonexamples/recipes/Alibaba-NLP_gte-reranker-modernbert-base/cpu/cpu/reranking_fp16_config.json/loader/tasktext-classificationreranking/eval56a6d0140cf6356659e2a7c1413286a774468d44, streaming true, shuffle false, samples 2, max candidates 10/export/compatibility{"transformers_attention":"eager"}/loader/tasktext-classificationreranking/quantnullmode=fp16andtask=reranking/eval56a6d0140cf6356659e2a7c1413286a774468d44, streaming true, shuffle false, samples 2, max candidates 10/export/compatibility{"transformers_attention":"eager"}Both recipes preserve
AutoModelForSequenceClassification,modernbert, opset 17, eager attention compatibility, INT32[1,1024]inputs, and logits output. The delta is reducibility-consistent with the charter. Recipe-free acceptance passed: publishedpipeline_tag=text-rankingresolves generically to canonicalrerankingwith no model-ID hard-coding. Generic source code is owned by dependency PR #1322 at parent3708969b731425b0c6d4b97920d1b5e6519bb013; this candidate has no code changes. Existing text-classification recipes andexamples/recipes/README.mdremain untouched.6. Analyze summary - component level and op level
Static rule analysis completed with
ANALYZE-PARTIAL-SUCCESSand exit code 1; this is compatibility analysis, not runtime execution. The unresolved overview gap is per-operator input-signature grouping: architecture regions are mapped through hierarchy scopes and tensor boundaries, while an ops-depth breakdown would be needed to resolve that finer grouping.Component-level summary
The mapped regions cover HTP-tagged attention and MLP scopes across all 22 encoder layers, an 11-node prediction-head region, and the scalar classifier tensor boundary. Embeddings, rotary embeddings, and mean pooling remain architecture-verified overview regions rather than operator-signature mappings.
Op-level summary
NvTensorRTRTX, OpenVINO GPU, and QNN GPU had rule-backed classifications but no runtime support on the test host; QNN's actionable partial types were
OP/ai.onnx/GatherandOP/ai.onnx/ReduceSum. CUDA, MIGraphX, TensorRT, and DML had no populated rule classification in this handoff. No unsupported operator type was reported.7. Reproduce commands
Set
$MODELto the pinned checkpoint directory and$OUTto a writable output directory. These are the tester-recorded build, perf, eval, and analyze command semantics with portable paths.