Skip to content

Add execution provider selection for builder capture - #2690

Merged
David Fan (jiafatom) merged 2 commits into
mainfrom
jiafa/fix-int4-model-builder-gpu-routing
Sep 26, 2026
Merged

David Fan (jiafatom) merged 2 commits into
mainfrom
jiafa/fix-int4-model-builder-gpu-routing

Conversation

@jiafatom

@jiafatom David Fan (jiafatom) commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • add --execution_provider to olive capture-onnx-graph for Model Builder and Mobius Builder exports
  • configure a separate Olive target system when the option is provided
  • keep --conversion_device scoped to exporters that need a conversion host device
  • preserve automatic CUDA routing for FP16/BF16 when the new option is omitted
  • add coverage for legacy INT4/FP16 defaults and CPU-host/CUDA-target exports through both builders

Why

Model Builder and Mobius Builder use the target execution provider to choose the generated graph format. In particular, a CUDA INT4 QMoE export uses CUDA-specific expert-weight layouts and FP16 I/O. The existing capture CLI derives that argument from Olive's target accelerator but cannot independently select a CUDA target for INT4.

This change uses Olive's existing host/target separation:

olive capture-onnx-graph \
  --use_model_builder \
  --precision int4 \
  --execution_provider CUDAExecutionProvider \
  -m <model> -o <output>

The export host uses the existing CPU default and does not need a GPU. The capture workflow does not run on or evaluate the target, so Olive skips the local supported-EP check, and ONNX Runtime GenAI CUDA weight prepacking has a CPU fallback.

--execution_provider currently requires --use_model_builder or --use_mobius_builder.

Testing

  • pytest test/cli/test_cli.py -k "capture_onnx_command_builder_execution_provider or capture_onnx_command_model_builder_accelerator or capture_onnx_command_use_mobius_builder or capture_onnx_command" -q (14 passed)
  • ruff check olive/cli/capture_onnx.py test/cli/test_cli.py
  • ruff format --check olive/cli/capture_onnx.py test/cli/test_cli.py

All tests were run in the jiafa-dev container.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

FP16/BF16 CPU routing regression remains unresolved; MobiusBuilder GPU coverage is also requested.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)
What changed in this PR

Fixes GPU routing for INT4 Model Builder and Mobius Builder exports.

Changes:

  • Routes explicit GPU conversions through CUDA.
  • Adds INT4 Model Builder GPU regression coverage.
File Description
test/​cli/​test_cli.py Adds INT4 Model Builder GPU configuration coverage.
olive/​cli/​capture_onnx.py Updates builder accelerator and execution-provider routing.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread olive/cli/capture_onnx.py Outdated
@jiafatom
David Fan (jiafatom) force-pushed the jiafa/fix-int4-model-builder-gpu-routing branch from c69d496 to 399a0ab Compare September 25, 2026 17:57
@jiafatom David Fan (jiafatom) changed the title Fix GPU routing for INT4 Model Builder exports Honor CUDA target for INT4 Model Builder exports Sep 25, 2026
@jiafatom
David Fan (jiafatom) force-pushed the jiafa/fix-int4-model-builder-gpu-routing branch from 399a0ab to 5b3f096 Compare September 25, 2026 18:05
@jiafatom David Fan (jiafatom) changed the title Honor CUDA target for INT4 Model Builder exports Add target execution provider for Model Builder capture Sep 25, 2026
@jiafatom
David Fan (jiafatom) force-pushed the jiafa/fix-int4-model-builder-gpu-routing branch from 5b3f096 to 0b246ee Compare September 25, 2026 18:08
Comment thread olive/cli/capture_onnx.py Outdated
@jiafatom
David Fan (jiafatom) force-pushed the jiafa/fix-int4-model-builder-gpu-routing branch from 0b246ee to 05e19bb Compare September 25, 2026 19:54
@jiafatom David Fan (jiafatom) changed the title Add target execution provider for Model Builder capture Add execution provider selection for builder capture Sep 25, 2026
@jiafatom
David Fan (jiafatom) enabled auto-merge (squash) September 25, 2026 21:06
Comment thread olive/cli/capture_onnx.py Outdated
@jiafatom
David Fan (jiafatom) merged commit a6a1080 into main Sep 26, 2026
12 checks passed
@jiafatom
David Fan (jiafatom) deleted the jiafa/fix-int4-model-builder-gpu-routing branch September 26, 2026 18:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants