Repository navigation
Conversation
审阅者指南(小型 PR 中默认折叠)审阅者指南该 PR 在截止长度计算周围添加了空结果保护:对于非空数据集保留现有行为;当预处理未生成有效 QA 对时发出警告,而不是调用下游处理。 空 QA 结果处理流程图flowchart TD
A[Preprocess chat records] --> B[save_result qa_res]
B --> C{qa_res is non-empty}
C -->|Yes| D["_execute_length_cdf_script()"]
C -->|No| E[logger.warning]
D --> F[Log processing success]
E --> F
文件级变更
可能相关的问题
提示和命令与 Sourcery 交互
自定义使用体验访问你的控制面板以:
获取帮助Original review guide in EnglishReviewer's guide (collapsed on small PRs)Reviewer's GuideThe PR adds an empty-result guard around cutoff length calculation, preserving existing behavior for non-empty datasets while warning instead of invoking downstream processing when preprocessing produces no valid QA pairs. Flow diagram for empty QA result handlingflowchart TD
A[Preprocess chat records] --> B[save_result qa_res]
B --> C{qa_res is non-empty}
C -->|Yes| D["_execute_length_cdf_script()"]
C -->|No| E[logger.warning]
D --> F[Log processing success]
E --> F
File-Level Changes
Possibly linked issues
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Empty-result cleaning can still trigger model inference before the new guard.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Fixes empty-dataset handling by skipping cutoff-length calculation and warning instead.
Changes:
- Preserves saving empty results.
- Skips length calculation for empty QA results.
- Reviewed cleaning path may still invoke inference before the guard.
| File | Summary |
|---|---|
weclone/data/qa_generator.py |
Adds empty-result handling, but enabled cleaning can still process empty inputs before the guard. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+150
to
+153
| if qa_res: | ||
| self._execute_length_cdf_script() | ||
| else: | ||
| logger.warning("No valid QA pairs generated; skipping cutoff_len calculation.") |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
cutoff_lencalculation when preprocessing produces no valid QA pairs.Why
When all messages are filtered or no QA pairs are matched,
length_cdf.pyis still invoked. LLaMA-Factory then fails withInstruction "train" corresponds to no data!, as seen in #92. There is no useful length statistic to compute for an empty dataset.Scope
This addresses the follow-on error shown in #92 without changing the filtering/matching logic that produced zero samples.
Sourcery 总结
错误修复:
Original summary in English
Summary by Sourcery
Bug Fixes: