Skip to content

introduces 64gb optimized qwen3.8-27b webgpu model - #653

Open
Prathik Rao (prathikr) wants to merge 1 commit into
mainfrom
prathikrao/qwen3.8-27b-webgpu-optimized
Open

Prathik Rao (prathikr) wants to merge 1 commit into
mainfrom
prathikrao/qwen3.8-27b-webgpu-optimized

Conversation

@prathikr

Copy link
Copy Markdown
Contributor

sets chunk_size and max_scheduled_tokens to 256 to maximize performance on Mac mini M4 Pro (64GB)

Copilot AI balanced review requested due to automatic review settings October 6, 2026 23:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

ONNX Runtime GenAI rejects runtime profiles for non-CUDA providers, preventing the generated WebGPU model from loading.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Adds an INT4 WebGPU recipe for Qwen3.8-27B, optimized for a 64 GB Mac configuration.

Changes:

  • Adds paged-attention WebGPU export configuration.
  • Adds memory-based runtime tuning.
  • Documents export and registers the recipe.
File Description
README.md Documents export and runtime settings.
Qwen-Qwen3.8-27B_webgpu_int4.json Defines model export and runtime configuration.
info.yml Registers the recipe metadata.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +74 to +75
"runtime_profiles": {
"type": "GenAIModelRuntimeProfiles",
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants