Skip to content

feat(laya): add native decision models and routing - #1614

Merged
yifeif-nv merged 7 commits into
NVIDIA:mainfrom
yifeif-nv:feat/laya-native
Oct 8, 2026
Merged

yifeif-nv merged 7 commits into
NVIDIA:mainfrom
yifeif-nv:feat/laya-native

Conversation

@yifeif-nv

@yifeif-nv yifeif-nv commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Background

Add the three checkpoints bundled in convaiinnovations/laya@7b928d828b7b0e022f929d9bd2e44165aa270148: English at the root, multilingual, and typed-decisions. The native implementation targets the released laya==0.3.20 request and response behavior.

Exit Criteria

All declared variants and model-card examples must pass original-reference accuracy checks, native inference must contain no Python/Torch dependency, matched complete Router calls must beat a validated compiled reference, and stable Community CI plus Internal CI must pass on the final head. The original per-option accuracy gates are unchanged. Derived response fields must follow the released SDK formulas, with exact four-decimal formatting of the validated native probabilities.

Implementation

One new family owns the ModernBERT encoder, decision/action heads, TensorRT plans, native BPE/NFC tokenization, question batching, calibration, response formatting, and routing. The build command selects one variant or a three-model router bundle. The C++ decide command and existing structured-decision Task execute both bundle kinds. No shared API, ABI, or bundle-framing change is required.

Change categories

  • Model or runtime behavior
  • Public API
  • ABI
  • Bundle or artifact format
  • Dependencies
  • Documentation only
  • CI or developer tooling

Validation

Commands and Results

  • python -m tools.model_ci validate: passed.
  • python -m pytest families/laya/tests/test_support.py -q: 11 passed.
  • python -m tools.legal_headers --check: no findings.
  • The BF16 GELU implementation matches the original CUDA activation over all 65,536 input bit patterns.
  • Combined architecture and support checks: 80 passed; mixed Clef/Laya support collection: 25 passed.
  • ruff check --config ruff.toml families/laya and git diff --check: passed.
  • python -m families.laya.tests.compare_record --checkpoint /artifacts/checkpoint --probe /src/build/laya_record_probe --output /artifacts/record-parity: each variant matched 2,099 tokenizer inputs and 18 complete records exactly, including Unicode normalization and long-conversation truncation.
  • python -m families.laya.tests.compare_routing --probe /src/build/laya_routing_probe --output /artifacts/routing-parity: 665 exact routing decisions and metadata comparisons passed.
  • python -m pytest families/laya/tests/test_e2e.py --e2e-model laya -q: 23 passed against the original CUDA SDK, using the qualified English, multilingual, typed-decisions, and three-model Router bundles. This covers the model-card examples, empty requests, single-option questions, long conversations, question batches, and repeated public responses. The native process loads no Python/Torch libraries.
  • Raw logits retain atol=0.125, rtol=0.015; per-option probabilities retain atol=0.002, rtol=0.01. Choices, usage, routing, and response structure match the reference; action probabilities retain their original gate. Entropy confidence and expected scores are checked with the SDK formulas on the validated native probabilities. Their numeric differences from original inference can reflect propagated probability error; the long-conversation confidence is 0.3635 versus 0.3695.
  • python -m families.laya.tests.compare_router --checkpoint /artifacts/checkpoint --bundle /artifacts/laya-router-qualified.bundle --probe /src/build/laya_task_probe --runtime-root /src/build --output /artifacts/router-behavior-qualified: all 22 explicit/automatic Router cases passed, with every variant exercised through the combined bundle.
  • Nine diagnostic corruption controls rejected wrong choice, entropy confidence, expected score, answer confidence, action probability, displayed probability, binary probability, usage, and routing.
  • Matched Router benchmarks after the precision fixes, with five warmups and 30 samples, measured native / validated torch.compile(max-autotune) p50: English email 2.44 / 5.41 ms, Hindi 2.45 / 3.71 ms, Spanish 1.89 / 2.60 ms, typed-decisions email 2.25 / 4.77 ms. Both sides measure the complete public Router call, excluding build/load/warmup.

Hardware, Environment, and Revisions

Head be695cddadc575a870f71a938d02aaa5405ebd94; checkpoint convaiinnovations/laya@7b928d828b7b0e022f929d9bd2e44165aa270148; SDK laya==0.3.20; Transformers 5.10.2; Torch 2.12.0+cu130; TensorRT 11.1.0.106. Local validation uses NVIDIA GB300, driver 580.105.08, and Ubuntu 24.04.4. Runtime graphs preserve the checkpoint's operand precision with explicit FP32 arithmetic boundaries.

Not Run / Remaining Gaps

Final-head Stable Community CI and Internal CI passed on be695cddadc575a870f71a938d02aaa5405ebd94. The optional 8192-token capacity and other GPU platforms have not been qualified. SDK callbacks and external integrations are outside the native Task interface. Local E2E reused previously built qualified bundles, except for the multilingual standalone bundle rebuilt from the current source; Internal CI also passed with freshly built bundles.

Contributor Self-Review

  • I have completed a self-review of this change.

Contributor self-review completed on head be695cddadc575a870f71a938d02aaa5405ebd94 against e6c674ebe329f109a6d4a5da20d08f472fe094c9. No blocking architecture or code findings remain; both required final-head CI lanes passed.

Notes For Future Readers

Start with model.py, then follow runtime/plugin.cpp into record, tokenizer, and routing behavior. The original SDK remains a build/reference dependency only. The router loads all three native engines; SDK callbacks and external integrations are outside the Task interface. Dev CI is informational; stable Community CI and Internal CI are the required readiness gates. Leave this PR unmerged.

Risk level

  • Low
  • Medium
  • High

The change is isolated to one new family. It adds a complete encoder, native tokenizer, and routing protocol; qualification currently covers the declared default capacities on GB300.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/TensorRT-Model-Connect/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: acd4d51e-de25-43c9-869b-435620a40b3e
📥 Commits

Reviewing files that changed from the base of the PR and between e6c674e and be695cd.

📒 Files selected for processing (49)
  • families/laya/NOTICE
  • families/laya/README.md
  • families/laya/__init__.py
  • families/laya/cli.json
  • families/laya/cli.py
  • families/laya/model.py
  • families/laya/requirements.txt
  • families/laya/routing.py
  • families/laya/runtime/CMakeLists.txt
  • families/laya/runtime/bpe_tokenizer.cpp
  • families/laya/runtime/cli.cpp
  • families/laya/runtime/plugin.cpp
  • families/laya/runtime/record.cpp
  • families/laya/runtime/record.h
  • families/laya/runtime/routing.cpp
  • families/laya/runtime/routing.h
  • families/laya/runtime/tokenizer.h
  • families/laya/runtime/unicode.h
  • families/laya/support.py
  • families/laya/tests/__init__.py
  • families/laya/tests/benchmark_compile.py
  • families/laya/tests/compare_plan.py
  • families/laya/tests/compare_record.py
  • families/laya/tests/compare_router.py
  • families/laya/tests/compare_routing.py
  • families/laya/tests/conftest.py
  • families/laya/tests/engine_probe.cpp
  • families/laya/tests/fixtures/batched.json
  • families/laya/tests/fixtures/conversation.json
  • families/laya/tests/fixtures/email.json
  • families/laya/tests/fixtures/empty.json
  • families/laya/tests/fixtures/hindi.json
  • families/laya/tests/fixtures/quickstart.json
  • families/laya/tests/fixtures/single-option.json
  • families/laya/tests/fixtures/spanish.json
  • families/laya/tests/fixtures/typed-email.json
  • families/laya/tests/manifests/laya-multilingual.json
  • families/laya/tests/manifests/laya-router.json
  • families/laya/tests/manifests/laya-typed-decisions.json
  • families/laya/tests/manifests/laya.json
  • families/laya/tests/record_probe.cpp
  • families/laya/tests/routing_probe.cpp
  • families/laya/tests/task_probe.cpp
  • families/laya/tests/test_e2e.py
  • families/laya/tests/test_gelu.py
  • families/laya/tests/test_routing.py
  • families/laya/tests/test_support.py
  • families/laya/tokenizer.py
  • website/data/hf-model-metadata.json

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


📝 Walkthrough

⚠️ A high-level summary could not be generated for this review. CodeRabbit will regenerate it on the next update, or you can request a refresh with @coderabbitai summary.

Walkthrough

Adds Laya structured-decision model support. The changes build TensorRT checkpoints, add native tokenization and inference with model routing, and provide fixtures and validation tools for the supported variants.

Changes

Laya structured-decision runtime

Layer / File(s) Summary
Checkpoint support and TensorRT plan
families/laya/cli.py, families/laya/cli.json, families/laya/model.py, families/laya/support.py, families/laya/tokenizer.py, families/laya/requirements.txt, families/laya/NOTICE, families/laya/__init__.py, website/data/hf-model-metadata.json, families/laya/README.md
Adds Laya model detection and build commands, validates build requests, and builds TensorRT plans for the supported variants. Adds tokenizer normalization data and checkpoint metadata. Documents variants, build and inference commands, limits, and supported inputs.
Native tokenization and record encoding
families/laya/runtime/tokenizer.h, families/laya/runtime/bpe_tokenizer.cpp, families/laya/runtime/record.h, families/laya/runtime/record.cpp, families/laya/runtime/unicode.h, families/laya/tests/record_probe.cpp, families/laya/tests/compare_record.py
Adds BPE tokenization and Unicode handling. Encodes structured-decision questions and formats choice, score, and noul answers. Adds a probe and comparisons against the released tokenizer and SDK.
Native inference and routing
families/laya/runtime/plugin.cpp, families/laya/runtime/routing.h, families/laya/runtime/routing.cpp, families/laya/runtime/cli.cpp, families/laya/runtime/CMakeLists.txt, families/laya/routing.py, families/laya/README.md
Adds the native inference pipeline and CLI. The router selects a variant using explicit model or task fields, typed-schema detection, language fields, or text analysis. The returned document includes the routing decision.
Fixtures and validation
families/laya/tests/*, families/laya/tests/fixtures/*, families/laya/tests/manifests/*, families/laya/runtime/CMakeLists.txt, families/laya/README.md
Adds manifests and fixtures for English, multilingual, typed-decision, batched, and edge-case inputs. Adds native probes and tests that compare plan outputs, tokenization, routing, compiled and eager results, and end-to-end responses.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant FamilyFactory
  participant Router
  participant Routing
  participant Pipeline
  participant TensorRTEngine
  Caller->>FamilyFactory: Load Laya bundle
  FamilyFactory->>Router: Create router when router.json is present
  Router->>Routing: Route document
  Routing-->>Router: Return selected model and routing decision
  Router->>Pipeline: Decide using selected model
  Pipeline->>TensorRTEngine: Run padded input tensors
  TensorRTEngine-->>Pipeline: Return logits and action logits
  Pipeline-->>Caller: Return answers, scores, usage, and routing decision
Loading

Merge Risk: ⚪ Minimal · up to be695

This change adds new, self-contained Laya model support and an optional router under families/laya. No concrete defect was found in the reviewed code. Final-head CI is reported as pending and should complete normally before merge.

🚥 Pre-merge checks | ✅ 6 | ❌ 2 | ❓ 1

❌ Failed checks (2 warnings, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 21.58% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 190 functions across 30 files. (19 skippe… Write docstrings for the functions missing them to satisfy the coverage threshold.
Benchmark Validation Integrity ⚠️ Warning The benchmark paths do not measure equivalent regions. families/laya/tests/benchmark_compile.py:53-55,101-107 times Router.predict on an already parsed Python record and synchronizes CUDA before a… Define one timing contract and use it for both implementations. Run the same logical record against the same topology, either a single variant on both sides or a three-variant router bundle on both sides. Align request parsing and response …
Shared Change Blast Radius ❓ Inconclusive The diff is family-local except for website/data/hf-model-metadata.json, a shared checkpoint catalog entry. Repository evidence says this file is documentation metadata only and that no CI job reads… Provide the complete pull request description or an explicit rationale for the shared catalog entry. Identify its consumers, document any behavior or compatibility impact, list validation for those consumers, and explain why the entry must …
✅ Passed checks (6 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Family Ownership Boundary ✅ Passed No cross-family ownership dependency is introduced. The changed implementation and validation assets reside under families/laya; Python imports use families.laya or shared core packages, and the n…
Shared Semantic Neutrality ✅ Passed PASS. The only changed path outside families/laya is website/data/hf-model-metadata.json. The change adds one documentation snapshot entry for the pinned Laya checkpoint. The file states that it i…
Title check ✅ Passed The title clearly identifies the addition of native Laya decision models and routing, which matches the primary changes.
Description check ✅ Passed The description covers the required background, exit criteria, implementation, change categories, validation results, environment, remaining gaps, self-review, notes, and risk level. It provides suffi…
Full details: Docstring Coverage

Explanation

Docstring coverage is 21.58% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 190 functions across 30 files. (19 skipped: 19 unsupported.)

Full details: Benchmark Validation Integrity

Explanation

The benchmark paths do not measure equivalent regions. families/laya/tests/benchmark_compile.py:53-55,101-107 times Router.predict on an already parsed Python record and synchronizes CUDA before and after the call. The native path families/laya/tests/task_probe.cpp:54-59,66-68 times IStructuredDecision::decide; families/laya/runtime/plugin.cpp:63-65,127-131 parses the document and serializes the response inside that timed call. Thus native timing includes per-call JSON parsing and response serialization that the Python timing excludes. The router topology is also not paired by the harness: benchmark_compile.py:23-25,47-48 benchmarks one attached variant, while compare_router.py:71-81 validates a native router but does not time it.

Resolution

Define one timing contract and use it for both implementations. Run the same logical record against the same topology, either a single variant on both sides or a three-variant router bundle on both sides. Align request parsing and response serialization: include both in both measurements, or exclude both with equivalent adapters. Keep CUDA/device synchronization around the same region. Use identical warmup and iteration counts, retain raw per-iteration samples, and compute p50/p95 from both paths under the same aggregation and task-unit definition. Record both measurements and the selected bundle identity in one comparison artifact.

Full details: Shared Change Blast Radius

Explanation

The diff is family-local except for website/data/hf-model-metadata.json, a shared checkpoint catalog entry. Repository evidence says this file is documentation metadata only and that no CI job reads it; no other repository consumer references the file. The repository also states that family code should remain isolated and shared code should be limited to model-agnostic contracts. The visible pull request text reports validation and says there is no shared API, ABI, or bundle-framing change, but the authored description is truncated. Therefore the required description evidence for the catalog's model-agnostic need, affected consumers, compatibility impact, and reason it cannot remain family-owned is unavailable.

Resolution

Provide the complete pull request description or an explicit rationale for the shared catalog entry. Identify its consumers, document any behavior or compatibility impact, list validation for those consumers, and explain why the entry must be maintained outside families/laya.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@yifeif-nv yifeif-nv added the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
@github-actions github-actions Bot removed the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
@yifeif-nv yifeif-nv added the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
@github-actions github-actions Bot removed the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
Add the English, multilingual, and typed-decisions checkpoints with an
optional native router. Keep model construction, tokenization, calibration,
runtime orchestration, and original-SDK validation inside the Laya family.

Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
Keep the original logit, probability, selected-choice, and action gates. Check derived fields with the released SDK formulas on the validated native probabilities, including exact rounding and response metadata. Require repeated public responses to match.

Signed-off-by: yifeif <277870278+yifeif-nv@users.noreply.github.com>
@yifeif-nv
yifeif-nv marked this pull request as ready for review October 8, 2026 17:27
@yifeif-nv yifeif-nv added the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
@github-actions github-actions Bot removed the run-internal-ci Maintainer-approved dispatch to internal CI label Oct 8, 2026
@yifeif-nv
yifeif-nv merged commit 0ab8320 into NVIDIA:main Oct 8, 2026
54 of 68 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant