Skip to content

Add multi-component export support for CLM and KEV - #749

Open
apsonawane wants to merge 5 commits into
mainfrom
asonawane/non-generative-export
Open

apsonawane wants to merge 5 commits into
mainfrom
asonawane/non-generative-export

Conversation

@apsonawane

@apsonawane apsonawane commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add generic multi-component model packaging to Mobius and implement model-specific export support for CLM-v0.1-8B and KEV-4B.

Changes

  • Add generic named components with neutral backbone/encoder/head roles and nested resolution.
  • Add CLM structured rendering, no-cache Qwen3 backbone export, input/output-normalized projection heads, scaled-cosine scoring, stable grouped softmax, and checkpoint validation.
  • Add KEV exact request encoding, no-cache Qwen3.5 backbone export, pointer Q/K scoring, typed answer mapping, and checkpoint validation.
  • Require every supplied checkpoint weight to bind and every package initializer to receive data.
  • Make grouped softmax mask multiplication safe for FP16/BF16 exports.
  • Add synthetic and opt-in full-checkpoint parity tests.
  • Document cache-free scoring interfaces and normalization contracts.
  • Apply repository Ruff formatting and lint fixes.

Commits

  • b9c45ca6 Add multi-component decision model exports
  • e619d3ba Test and document decision model packages
  • 89e98d5c Harden CLM and KEV export contracts

Validation

  • Ruff checks passed
  • 31 tests passed
  • 3 opt-in checkpoint tests skipped by default
  • Regressions cover no-cache backbone interfaces, strict weight application, CLM normalization, and grouped-softmax dtype handling

Notes

  • No model weights or generated model packages are included.
  • Full-checkpoint parity tests require explicitly configured local artifacts.
  • CLM and KEV are exported as non-generative multi-component packages.

Introduce generic named-component task packaging and add CLM and KEV model adapters, checkpoint validation, preprocessing, heads, and export tasks.
Cover generic component manifests, CLM and KEV graph contracts, synthetic parity, and opt-in checkpoint parity while documenting package metadata and catalog layout.
@CLAassistant

CLAassistant commented Sep 29, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

Comment thread src/mobius/models/decision.py Fixed
Comment thread src/mobius/models/decision.py Fixed
Comment thread src/mobius/models/decision_test.py Fixed
Comment thread src/mobius/models/decision_test.py Fixed
Comment thread src/mobius/models/decision_test.py Fixed
Comment thread src/mobius/models/decision_test.py Fixed
Comment thread src/mobius/models/decision_test.py Fixed
Comment thread src/mobius/tasks/_decision.py Fixed
Comment thread tests/clm_parity_test.py Fixed
Comment thread tests/kev_parity_test.py Fixed
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Performance Comparison

Comparing 88fd6a1f → 55c576a0

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0% ⚪
bert (feature-extraction) num_nodes 68 68 +0.0% ⚪
falcon model_size_bytes 364 KB 364 KB +0.0% ⚪
falcon num_nodes 66 66 +0.0% ⚪
gemma2 model_size_bytes 428 KB 428 KB +0.0% ⚪
gemma2 num_nodes 105 105 +0.0% ⚪
gpt2 model_size_bytes 324 KB 324 KB +0.0% ⚪
gpt2 num_nodes 54 54 +0.0% ⚪
llama model_size_bytes 425 KB 425 KB +0.0% ⚪
llama num_nodes 60 60 +0.0% ⚪
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
llama (static-cache) num_nodes 56 56 +0.0% ⚪
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0% ⚪
mamba (ssm-text-generation) num_nodes 94 94 +0.0% ⚪
phi3 model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 num_nodes 58 58 +0.0% ⚪
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 (static-cache) num_nodes 54 54 +0.0% ⚪
qwen2 model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 num_nodes 60 60 +0.0% ⚪
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 (static-cache) num_nodes 56 56 +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0% ⚪
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0% ⚪
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) num_nodes 450 450 +0.0% ⚪
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0% ⚪
t5 (seq2seq) num_nodes 176 176 +0.0% ⚪
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0% ⚪
whisper (speech-to-text) num_nodes 128 128 +0.0% ⚪

No performance regressions.

Comment thread src/mobius/tasks/_base.py Fixed
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 88fd6a1f → 55c576a0

Model Sub-model Changes Status
bert (feature-extraction) model 0 ⚪
falcon model 0 ⚪
gemma2 model 0 ⚪
gemma4 (gemma4) decoder 0 ⚪
gemma4 (gemma4) embedding 0 ⚪
gemma4 (gemma4) vision_encoder 0 ⚪
gemma4_text model 0 ⚪
gpt2 model 0 ⚪
llama model 0 ⚪
llama (static-cache) model 0 ⚪
mamba (ssm-text-generation) model 0 ⚪
phi3 model 0 ⚪
phi3 (static-cache) model 0 ⚪
qwen model 0 ⚪
qwen (static-cache) model 0 ⚪
qwen2 model 0 ⚪
qwen2 (static-cache) model 0 ⚪
qwen2_moe model 0 ⚪
qwen2_moe (static-cache) model 0 ⚪
qwen3 model 0 ⚪
qwen3 (static-cache) model 0 ⚪
qwen3_5_moe (hybrid-text-generation) model 0 ⚪
qwen3_5_text (hybrid-text-generation) model 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) decoder 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) embedding 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 0 ⚪
qwen3_moe model 0 ⚪
qwen3_moe (static-cache) model 0 ⚪
qwen3_next (hybrid-text-generation) model 0 ⚪
t5 (seq2seq) decoder 0 ⚪
t5 (seq2seq) encoder 0 ⚪
whisper (speech-to-text) decoder 0 ⚪
whisper (speech-to-text) encoder 0 ⚪

No architecture changes detected. ✅


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Export contracts, weight validation, CLM normalization, and reduced-precision grouped softmax have unresolved correctness issues.

Review effort: Balanced
Findings: 2 High severity · 3 Medium severity

Open (5)
What changed in this PR

Adds generic multi-component packaging and specialized non-generative CLM/KEV exports.

Changes:

  • Introduces named components with neutral roles and nested resolution.
  • Adds CLM and KEV preprocessing, heads, scoring, validation, and packaging.
  • Adds parity tests and user documentation.
File Description
tests/​kev_parity_test.py Adds KEV synthetic and checkpoint parity tests.
tests/​clm_parity_test.py Adds CLM numerical and vLLM parity tests.
src/​mobius/​tasks/​_multi_component_test.py Tests generic component packaging.
src/​mobius/​tasks/​_decision.py Defines CLM and KEV export tasks.
src/​mobius/​tasks/​_base.py Adds component roles and generic task support.
src/​mobius/​tasks/​__init__.py Exports and registers new tasks.
src/​mobius/​models/​decision.py Implements decision models and adapters.
src/​mobius/​models/​decision_test.py Tests decision-model contracts and helpers.
src/​mobius/​models/​__init__.py Exports decision-model APIs.
src/​mobius/​__init__.py Exposes new public APIs.
docs/​model-catalog.md Documents CLM and KEV exports.
docs/​api/​model_package.md Documents generic multi-component tasks.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/mobius/models/decision.py
Comment thread src/mobius/models/decision.py
Comment thread src/mobius/tasks/_decision.py
Comment thread src/mobius/tasks/_decision.py
Comment thread src/mobius/tasks/_decision.py
Require complete checkpoint weight application, normalize CLM inputs, make grouped softmax dtype-safe, export cache-free scoring backbones, and apply repository formatting and lint fixes with focused regressions.
Comment thread src/mobius/models/decision_test.py Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

FP16 normalization, weight-attribution, component-role, and validation issues remain unresolved.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)
Resolved since last review (5)
Previously missed (3)

In code that hasn't changed since last review

Medium severity Validate aligned inputs and answer types in the public mapper

src/​mobius/​models/​decision.py:811

This public answer mapper does not validate its aligned inputs: unequal lengths are silently truncated by zip, extra probabilities can index past keys, an empty choice reaches max() on an empty range, and an unknown type falls through as a score. Add the same non-empty/equal-length and type validation used by clm_answer, plus the two-value invariant for noul.

Medium severity Validate components against merged model roles

src/​mobius/​tasks/​_base.py:334

Validation reads only components.roles(), while __init_subclass__ explicitly merges subclass model_roles. A task using plain component paths plus explicit roles therefore still fails here; in a mixed spec, an unclassified extra component can instead pass the count check, get built, and default to decoder optimization. Validate all component names against the merged self.model_roles so both supported declaration styles enforce the same invariant.

Medium severity Fail parity validation for incomplete explicitly configured exports

tests/​kev_parity_test.py:325

Once MOBIUS_DECISION_EXPORT_ROOT is explicitly set, an incomplete package should fail this parity test rather than skip it. As written, a broken export reports only a skip, so the opt-in validation can appear clean without exercising either exported graph.

Comment thread src/mobius/_model_package.py
Comment thread src/mobius/models/decision.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Checkpoint mappings can override validated base weights, and several advertised provenance and dtype guarantees remain unenforced or untested.

Review effort: Balanced
Findings: None

Resolved since last review (2)
Previously missed (6)

In code that hasn't changed since last review

Medium severity Public mapper lacks input cardinality validation

src/​mobius/​models/​decision.py:818

Unlike clm_answer, this public mapper does not validate key/probability cardinality. A short noul distribution raises IndexError, while mismatched choice inputs can truncate the returned distribution or index past keys. Reject mismatched/empty inputs and require exactly two noul probabilities.

Medium severity Provenance marks mutable revisions as reproducible

src/​mobius/​models/​decision.py:1062

Any nonblank value, including a movable branch such as main, is stamped with reproducible=true by ModelProvenance.as_metadata(). This contradicts the documented <sha> contract and can make an export impossible to reproduce later. Require and normalize an immutable 40-hex revision before stamping it.

Medium severity Head mapping can overwrite validated base checkpoint weights

src/​mobius/​models/​decision.py:1086

map_clm_checkpoint() also accepts top-level encoder.* and flattened head tensors, and this update runs after the separately supplied base mapping. Consequently, an otherwise valid head checkpoint can silently override base weights or the validated nested head tensors. Restrict this production path to the fields validated by validate_clm_checkpoint().

Medium severity Head mapper imports unvalidated backbone tensors

src/​mobius/​models/​decision.py:1138

map_kev_checkpoint() imports top-level backbone.*/model.* tensors as well as the validated pointer head. Because this update follows the merged-base mapping, such entries in head_checkpoint silently replace the caller's merged base. Limit the production helper to the validated head state dict.

Medium severity build() ignores overridden model_roles

src/​mobius/​tasks/​_base.py:327

__init_subclass__ merges an explicit model_roles override into the derived role map, but build() ignores that map and re-reads only roles embedded in ComponentConfig. A subclass using plain component paths plus model_roles is therefore rejected, and an explicit role override can disagree with the role used for validation. Use the derived self.model_roles values for the declared component names here.

Low severity Missing FP16/BF16 grouped-softmax regression coverage

src/​mobius/​models/​decision.py:920

The claimed FP16/BF16 grouped-softmax regression is not exercised: both parity fixtures build FLOAT graphs, and the reduced-precision unit test only inspects projection-head casts. Add FLOAT16 and BFLOAT16 scorer/pointer coverage with nontrivial groups, asserting finite outputs and per-group sums of one, so mask/sentinel/Exp dtype regressions cannot pass CI.

@apsonawane
apsonawane marked this pull request as ready for review September 30, 2026 19:42
@apsonawane
apsonawane requested a review from a team September 30, 2026 19:42
@titaiwangms

titaiwangms commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

@xiaoyu-work Can you take a look? I think there might be some overlapping work here that you already done (multi-component)?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants