Skip to content

fix(data): supervise historical Qwen thinking turns - #933

Open
cheater6666 wants to merge 1 commit into
sgl-project:mainfrom
cheater6666:fix/qwen-history-loss-mask
Open

cheater6666 wants to merge 1 commit into
sgl-project:mainfrom
cheater6666:fix/qwen-history-loss-mask

Conversation

@cheater6666

Copy link
Copy Markdown

Motivation

Qwen3.5 and Qwen3-Next-Thinking render historical assistant turns without the <think> scaffold used by the final assistant turn. The parser only matched the longer configured header, silently masking all historical assistant outputs from training.

Modifications

  • Allow chat templates to declare alternate assistant headers.
  • Match headers longest-first so the final-turn thinking scaffold remains excluded.
  • Register historical headers for Qwen3.5 and Qwen3-Next-Thinking.
  • Add a focused parser regression test and refresh the real-tokenizer references.

Related Issues

Fixes #912

Accuracy Test

Input IDs are unchanged. All changed loss-mask values are 0 -> 1 in historical assistant spans.

  • Qwen3.5 supervised tokens: 146 -> 257
  • Qwen3-Next-Thinking supervised tokens: 125 -> 219
python -m unittest \
  tests.test_data.test_assistant_header_alternates \
  tests.test_data.test_parsers.TestTemplatePreprocessing.test_qwen35_instruct \
  tests.test_data.test_parsers.TestTemplatePreprocessing.test_qwen3_next_thinking -v

Result: 3 tests passed.

Benchmark & Profiling

Not applicable; this changes preprocessing correctness, not runtime inference performance.

Checklist

  • Ran the repository pre-commit hooks on all changed files.
  • Added a unit test and updated real-tokenizer regression references.
  • Documentation changes are not needed for this bug fix.
  • Performance and model accuracy benchmarks are not applicable.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] qwen3.5 / qwen3-next-thinking loss mask only supervises the last assistant turn

1 participant