Skip to content

无法复现 LongVT-7B-SFT 在 VideoMME 上的结果 #25

Description

@Physicsmile

感谢你们出色的工作和详细的 FAQ。我按照 README/FAQ 操作,用官方的开源权重在VideoMME上进行测评,但始终比论文报告(表 2,“密集帧采样”,含字幕)中 LongVT-7B-SFT 在 VideoMME 上的 64.9% 低了约 11 %。我已经排除了 FAQ 中列出的常见问题,因此我将报告完整设置和已检查的内容,以防我遗漏了某个配置细节。

结果

指标 数值
确定性准确率(严格字母匹配) 0.3974
使用 LLM 评判的准确率(Qwen2.5-72B-Instruct) 0.5367
格式得分 0.7256
论文表 2(LongVT-7B-SFT,VideoMME 含字幕,密集采样) 64.9

按时长划分(评判准确率):短视频 65.4% / 中等视频 53.2% / 长视频 42.3%。短视频基本复现;全部差距都出现在中等视频,尤其是长视频中。这个结果在我约 7 次运行中都很稳定(得分都在 52–54% 之间)。

环境 / 确切配置

  • 模型:已发布 LongVT-7B-SFT 检查点(SFT 版本,基础模型为 Qwen2.5-VL-7B-Instruct)
  • 任务:videomme_w_subtitle_reward_tool(已验证字幕确实被注入:2700 个样本中有 2232 个包含真实字幕文本;其余样本的数据集中没有 .srt)
  • 推理(vLLM serve):vLLM 0.12.0,--chat-template examples/eval/tool_call_qwen2_5_vl.jinja,--tool-call-parser hermes,--enable-auto-tool-choice,TP=4
  • 客户端(async_openai):fps=1, max_frames=768, max_pixels=50176, is_qwen3_vl=False, num_cpus=1, timeout=12000
  • 解码:贪心(do_sample=False,temperature=0)
  • max_new_tokens:49152(依据 FAQ Q2.5)
  • 运行规模:完整 2700 个样本,accelerate --num_processes=4
  • 评判模型:Qwen2.5-72B-Instruct,本地独立端口提供服务,USE_LLM_JUDGE=True

FAQ 中已检查并确认不是原因的问题

  • ✅ 任务名称是 videomme_w_subtitle_reward_tool(已确认提示中包含了字幕)。
  • ✅ IS_QWEN3_VL=False。
  • ✅ vLLM 0.12.0 + 提供的聊天模板 + hermes 解析器。
  • ✅ max_new_tokens=49152(只有 2.2% 的响应缺少 <answer> 标签或被截断)。
  • ✅ 完整的 2700 个样本(未使用 --limit)。
  • ✅ MCP crop_video 路径处理正常(使用干净的绝对路径;已配置前导斜杠的回退方案)。
  • ✅ hermes 工具解析器的 JSON 错误极少:在长达 8 小时的运行中,仅 8 / 2455 个工具调用样本(0.3%)出现 JSONDecodeError。
  • ✅ 评判模型计分偏低不是问题所在:72B 评判模型已经几乎恢复了所有“用选项文本而非字母回答”的情况;只有约 10 个“答案字面等于正确选项文本但得分仍为 0”的样本残留。

问题

  1. 评判模型:报告中的 64.9 分是使用 Qwen3-235B-A22B 还是 Qwen2.5-72B-Instruct 生成的?对于 SFT 模型,评判模型的选择大约会使 VideoMME 分数移动多少个百分点?(我的分数差距主要集中在中等/长视频上,而这也正是开放式/语义密集型视频。)
  2. 帧数(密集采样 = {512, 768} 中的最佳值):对于 LongVT-7B-SFT 而言,64.9 分是基于 512 帧还是 768 帧?512 帧与 768 帧之间的分数差异有多大?(我目前只有 768 帧的结果 = 53.7;512 帧的运行正在进行中。)
  3. 长视频上的工具行为:在我的运行中,长视频上的 crop_video 表现退化——78% 的裁剪起始时间 <1 秒(典型窗口为 [0, ~32秒],而视频长度为 30–60 分钟),单次裁剪且无精细化,工具没有带来提升(无裁剪准确率 0.469 > 起始锚定裁剪准确率 0.408)。这是已发布 SFT 检查点的预期行为吗?还是说它应该能进行时间定位?另外,对于 Qwen2.5-VL(is_qwen3_vl=False)路径,全局帧是以无时间标记的独立 image_url 形式发送的,每帧的时间戳标记(<X seconds>)仅在 Qwen3-VL 路径中添加——那么 SFT 模型在评估时是否依赖任何时间戳/帧时间信号来确定裁剪时间?
  4. 我们是否还遗漏了实现 64.9 分的其他设置细节(系统提示变体、评判模型投票次数、像素预算等)?

所使用的系统提示(来自 lmms_eval_tasks/videomme/utils.py):

"You are a helpful assistant. When the user asks a question, your response must include two parts: first, the reasoning process enclosed in <think>...</think> tags, then the final answer enclosed in <answer>...</answer> tags. ..."

另外附上完整的推理 JSON 文件

20260204_234358_samples_videomme_reward_tool.json
20260622_102318_results.json

谢谢!

Activity

  1. self-assigned this
    on Jun 24, 2026
  2. mwxely commented on Jun 24, 2026

    @mwxely
    Collaborator

    Hi @Physicsmile , thank you for the careful and well-instrumented reproduction. You got essentially all of the configuration right (I will confirm which below), so this comes down mostly to the eval judge (whether it is on, and which model) plus the frame protocol, with a smaller answer-format tail.

    First, line up the metrics

    These three numbers are not the same quantity, which is the first thing to disentangle:

    • 32.78 (videomme_perception_score) is a strict single-letter exact match (A–D). It is intentionally harsh and is not the headline metric, so do not compare it to 64.9.
    • 39.74 (acc_score in your results.json) and the paper's 64.9 are the same metric (acc_score): the repo's hybrid scorer = deterministic validators (relax_exact_match + math_verify) with an LLM-judge fallback. The repo computes both acc_score and perception_score for every run; the reported 64.9 is the acc_score (LLM-judge) path that the paper's evaluation describes (Sec 4.2 / Fig 6), not the strict perception_score. So 39.74 vs 64.9 is a like-for-like acc_score comparison, and the question is what suppresses it.
    • Your 53.67 is acc_score with your Qwen2.5-72B judge active.

    Where your points are: walk the three numbers

    Your own three numbers decompose the gap cleanly, and the single biggest lever is the judge:

    • 39.74 → 53.67 (+~14 pts): turn the judge on. The acc_score = 0.3974 in your results.json is consistent with the LLM judge being inactive in that run. The repo ships USE_LLM_JUDGE = False by default (custom_rewards/lmms_lab_recipe.py), so without it the score is essentially the deterministic-match rate, and your own 53.67 with the Qwen2.5-72B judge active confirms that activating it recovers ~14 points. Step one is therefore just: set USE_LLM_JUDGE=True and make sure the judge endpoint is reachable. (An unreachable judge silently scores every non-exact answer as 0, which by itself looks like a large gap.)
    • 53.67 → 64.9: judge model + frames + a small format tail. For the reported numbers we used Qwen3-235B-A22B as the eval judge, not 72B; this is the judge listed in the README reproduce table, and it is the judge server we ran our evaluations against at the time. Moving 72B → 235B closes part of the remaining gap; the rest is the frame protocol (Q2) and a residual answer-format tail (below). A gpt-oss-120b / deepseek-v4-class judge is an equally good modern choice if you would rather not stand up the 235B.

    One precision note so there are no surprises: the released scorer uses a binary {1, 0} judge (1 for a fully correct answer, 0 otherwise), called only as a fallback when the deterministic check returns 0. The paper's Figure 6 ("Evaluation Prompt for LLM-as-a-Judge") describes a three-way {1, 0.5, 0} evaluation judge (the same 1/0.5/0 scale is also the RL accuracy reward in Sec 4.2). On an all-MCQ benchmark like VideoMME the practical difference between binary and three-way is small (partial credit rarely triggers against a single-letter target), but the public code path is the stricter binary one.

    The residual format / extraction tail

    After the judge is on, the remaining recoverable chunk is answer-format. Your format_score = 0.7256 measures answer-format compliance; samples that never emit a clean <answer> (or \boxed{}) tend to score 0 on the deterministic path, and although the judge fallback can rescue some short responses, for long multi-turn tool answers (>500 chars) that whole-string fallback is gated out. Your reported "2.2% missing <answer>" and "0.3% JSON decode errors" feed exactly this. I went through the 100-sample slice you shared: every clean single-letter <answer> that matched the gold was already scored correctly by the deterministic matcher (no judge needed), and the acc_score=0 cases were either a genuinely wrong letter or a missing <answer>. So the judge's job on this MCQ benchmark is specifically to recover the missing-<answer> cases whose reasoning reached the right option, which is why turning it on helps but does not move everything.

    Q1. Judge model

    For the reported numbers the eval judge was Qwen3-235B-A22B (the judge listed in the README reproduce table), not 72B. Qwen2.5-72B-Instruct, which you used, is what we use for the RL training reward; as a 72B → 235B eval-judge swap it accounts for part of your 53.67 → 64.9 gap. We did not run a controlled ablation isolating the exact percentage-point impact, so I will not put a hard number on it, but on top of activating the judge it is a real contributor here. A gpt-oss-120b / deepseek-v4-class judge is an equally good modern alternative.

    Q2. Frames (512 vs 768)

    This is the config delta I would focus on, and the README's reproduce table is a little too terse about it. The paper's actual protocol (Experimental Setup, "Dense Frame Sampling" regime) is: 512 or 768 uniformly-sampled frames, and we report the better of the two. All three stages were trained at 512 frames (training-hyperparameters table), and the SFT checkpoint is highly frame-count-sensitive: the same SFT model scores only 12.5 under sparse 64-frame sampling vs 64.9 dense, a ~52-point swing.

    You ran 768 only. Because the SFT checkpoint is train-aligned at 512 and so frame-sensitive, 768 may simply be the worse of the two for it. Please also run max_frames=512 and take the better of {512, 768}; that is what the reported number reflects. (For the RFT checkpoint, 512 is in fact our documented default, which is probably the source of any "512" guidance you have seen.) I will fix the FAQ to state the better-of-two protocol explicitly rather than a flat 768.

    Q3. Tool behavior on long videos

    A few things are simultaneously true here, and I would rather be straight than defensive:

    1. LongVT-7B-SFT is the cold-start checkpoint, stage one of a three-stage SFT → RL → RFT pipeline, and the weakest released stage (VideoMME dense: SFT 64.9 < RL 66.1 < RFT 67.0). The paper is explicit that disciplined tool use (when to crop, how long a window) is what RL (with an IoU reward that curbs span inflation) and RFT stabilize, and that SFT alone yields only modest long-video gains. At the SFT stage the tool fires but is not yet fully disciplined, so on VideoMME its net accuracy contribution is small and the 64.9 is carried mostly by dense-frame coverage; the tool-use payoff shows up after RL/RFT. So a smaller tool contribution at the SFT stage is consistent with the paper's own staging, where RL and RFT are what make tool use pay off. That does not excuse the SFT number (64.9 is published and you deserve to hit it), but the headline tool-use behavior lives in RFT.
    2. Mechanism: crop_video(video_path, start_time, end_time) takes absolute timestamps in seconds that the model itself predicts from its temporal grounding; there is no external timestamp or frame-time signal injected (the chat template emits only <|video_pad|>, no per-frame time tokens). The tool enforces only start < end and in-bounds timestamps; it does not impose a minimum window duration. On your "78% of crops start < 1 s" observation, I would be careful reading that as a failure mode, because early crop starts are common even in the healthy model: in the Questions about the evaluation. #19 RFT reference log (below), about 50% of crop start times are < 1 s (median 2 s, mean ~91 s), so a model frequently cropping from near the beginning is partly expected, not by itself a sign of breakage. Your SFT number being higher (78%) is plausible for a cold-start checkpoint with less refined windowing, but I would not over-attribute the accuracy gap to it. I no longer have the LongVT-7B-SFT JSONL (it was an intermediate checkpoint and I have since left the previous company), so I cannot reproduce your exact 78% on my side.
    3. For a healthy reference, compare against the eval log I posted in Questions about the evaluation. #19 (gdrive): it is the RFT model on VideoMME w/sub + tool, it scores 66.96% (1,808 / 2,700), tools fire in 2,418 / 2,700 samples, and the <tool_call>crop_video {...} traces show the timestamp-driven cropping the headline numbers come from. That makes a useful side-by-side: if your SFT tool calls look structurally different (for example crops that overshoot the video duration, or malformed <tool_call> JSON), grepping your JSONL for crop_video errors against that reference will surface it. But as the metric decomposition above shows, the recoverable accuracy for the SFT checkpoint is in the judge and frame settings, not primarily in tool behavior.

    Q4. Other settings (system prompt, judge voting, pixel budget)

    • System prompt / tool framing: this is governed by the provided chat template (examples/eval/tool_call_qwen2_5_vl.jinja) that vLLM applies at serving time, not by a hidden client-side knob. When the tool schema is passed, that template emits the # Tools system block (rendering your examples/video_tools/mcp_server.py tool definitions) and the hermes <tool_call> format, so as long as you serve with --chat-template tool_call_qwen2_5_vl.jinja and --tool-call-parser hermes (both are in examples/eval/run_eval.sh) your tool and system framing match ours. One detail worth knowing: the videomme_w_subtitle_reward_tool task delivers the question through doc_to_text and does not prepend the repo's longer SYSTEM_PROMPT (the explicit <think>...</think> then <answer>...</answer> instruction); but that is identical for our runs and yours, so it is not a source of divergence between us. There is no separate async_openai version you need to match here.
    • Judge voting: there is no voting / self-consistency (no n>1, no majority vote). The scorer makes a single judge call with max_tokens=5 (plus at most one fallback call on the raw answer string).
    • Pixel budget: max_pixels=50176 (= 224×224) is exactly right; it is the eval default and matches the training config (video_max_pixels: 50176), so train/test are aligned. No change needed.

    What you already got right

    vLLM 0.12.0 (yes, that is our recommended eval version, with the provided tool_call_qwen2_5_vl.jinja chat template and --tool-call-parser hermes; other vLLM versions have incompatible hermes_tool_parser behavior, so keep 0.12.0), the task name videomme_w_subtitle_reward_tool, is_qwen3_vl=False (correct, all three checkpoints are fine-tuned from Qwen2.5-VL-7B-Instruct), max_pixels=50176, and max_new_tokens=49152 are all correct. Your ~610 output tokens/sample means you are on average not truncating (truncation was the failure mode in #19, fixed by the 49152 override). One safeguard worth applying regardless: if you spot any individual responses cut off before the <answer> tag, raise max_new_tokens further (for example 65536) and re-run, so every question is answered in full, since a truncated response scores 0 regardless of the judge. tp=4 versus our dp=8 serving is a throughput choice and should not move the score.

    Suggested next steps (most actionable first)

    1. Turn the judge on (USE_LLM_JUDGE=True, reachable endpoint). Your acc_score=0.3974 is consistent with it being off; your own 53.67 with the 72B judge shows ~+14 pts from activating it alone. (An unreachable judge silently zeroes every non-exact answer.)
    2. Use the Qwen3-235B-A22B judge we used for the reported number (or a gpt-oss-120b / deepseek-v4-class judge); the 72B → 235B swap closes part of the remaining 53.67 → 64.9 gap.
    3. Run max_frames=512 as well and take the better of {512, 768}, matching the reported protocol; the SFT checkpoint is train-aligned at 512.
    4. Mop up the format tail. Grep acc_score=0 samples for a correct letter in the reasoning but no clean <answer> (your format_score is 0.7256); with the judge on, the remainder is mostly genuine misses.
    5. If your goal is the headline tool-use behavior, also evaluate LongVT-7B-RFT and compare its trace against the Questions about the evaluation. #19 reference log, but the SFT 64.9 should be recoverable via steps 1–4.

    Thanks again for the rigor; this is exactly the kind of report that improves the repo, and I will tighten the FAQ on the 512/768 "better-of-two" protocol and the judge/format notes. Happy to look at specific JSONL cases if you can share a handful.

  3. added theissue type on Jun 24, 2026
  4. Physicsmile commented on Jun 24, 2026

    @Physicsmile
    Author

    Thanks, this clears up almost everything. Two quick clarifications so we can isolate the remaining gap:

    • Judge was already on — my 53.67 is the acc_score path with USE_LLM_JUDGE=True + a reachable 72B endpoint (the 39.74 was the judge-off results.json), so I'm past step 1.
    • I ran both frame counts — 768→53.67, 512→54.74, best-of = 54.7 (the long-video bucket is what gains at 512). So frames is handled too.

    I'll try more judge model(gpt-oss-120b / deepseek-v4-class) and re-score the deterministic-0 set without the >500-char gate.

    Thank you again for your patient reply!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

good first issueGood for newcomersquestionFurther information is requested

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions