Repository navigation
无法复现 LongVT-7B-SFT 在 VideoMME 上的结果 #25
Description
Activity
Hi @Physicsmile , thank you for the careful and well-instrumented reproduction. You got essentially all of the configuration right (I will confirm which below), so this comes down mostly to the eval judge (whether it is on, and which model) plus the frame protocol, with a smaller answer-format tail.
First, line up the metrics
These three numbers are not the same quantity, which is the first thing to disentangle:
- 32.78 (
videomme_perception_score) is a strict single-letter exact match (A–D). It is intentionally harsh and is not the headline metric, so do not compare it to 64.9. - 39.74 (
acc_scorein yourresults.json) and the paper's 64.9 are the same metric (acc_score): the repo's hybrid scorer = deterministic validators (relax_exact_match+math_verify) with an LLM-judge fallback. The repo computes bothacc_scoreandperception_scorefor every run; the reported 64.9 is theacc_score(LLM-judge) path that the paper's evaluation describes (Sec 4.2 / Fig 6), not the strictperception_score. So 39.74 vs 64.9 is a like-for-likeacc_scorecomparison, and the question is what suppresses it. - Your 53.67 is
acc_scorewith your Qwen2.5-72B judge active.
Where your points are: walk the three numbers
Your own three numbers decompose the gap cleanly, and the single biggest lever is the judge:
- 39.74 → 53.67 (+~14 pts): turn the judge on. The
acc_score = 0.3974in yourresults.jsonis consistent with the LLM judge being inactive in that run. The repo shipsUSE_LLM_JUDGE = Falseby default (custom_rewards/lmms_lab_recipe.py), so without it the score is essentially the deterministic-match rate, and your own 53.67 with the Qwen2.5-72B judge active confirms that activating it recovers ~14 points. Step one is therefore just: setUSE_LLM_JUDGE=Trueand make sure the judge endpoint is reachable. (An unreachable judge silently scores every non-exact answer as 0, which by itself looks like a large gap.) - 53.67 → 64.9: judge model + frames + a small format tail. For the reported numbers we used
Qwen3-235B-A22Bas the eval judge, not 72B; this is the judge listed in the README reproduce table, and it is the judge server we ran our evaluations against at the time. Moving 72B → 235B closes part of the remaining gap; the rest is the frame protocol (Q2) and a residual answer-format tail (below). Agpt-oss-120b/deepseek-v4-class judge is an equally good modern choice if you would rather not stand up the 235B.
One precision note so there are no surprises: the released scorer uses a binary {1, 0} judge (1 for a fully correct answer, 0 otherwise), called only as a fallback when the deterministic check returns 0. The paper's Figure 6 ("Evaluation Prompt for LLM-as-a-Judge") describes a three-way {1, 0.5, 0} evaluation judge (the same 1/0.5/0 scale is also the RL accuracy reward in Sec 4.2). On an all-MCQ benchmark like VideoMME the practical difference between binary and three-way is small (partial credit rarely triggers against a single-letter target), but the public code path is the stricter binary one.
The residual format / extraction tail
After the judge is on, the remaining recoverable chunk is answer-format. Your
format_score = 0.7256measures answer-format compliance; samples that never emit a clean<answer>(or\boxed{}) tend to score 0 on the deterministic path, and although the judge fallback can rescue some short responses, for long multi-turn tool answers (>500 chars) that whole-string fallback is gated out. Your reported "2.2% missing<answer>" and "0.3% JSON decode errors" feed exactly this. I went through the 100-sample slice you shared: every clean single-letter<answer>that matched the gold was already scored correctly by the deterministic matcher (no judge needed), and theacc_score=0cases were either a genuinely wrong letter or a missing<answer>. So the judge's job on this MCQ benchmark is specifically to recover the missing-<answer>cases whose reasoning reached the right option, which is why turning it on helps but does not move everything.Q1. Judge model
For the reported numbers the eval judge was
Qwen3-235B-A22B(the judge listed in the README reproduce table), not 72B.Qwen2.5-72B-Instruct, which you used, is what we use for the RL training reward; as a 72B → 235B eval-judge swap it accounts for part of your 53.67 → 64.9 gap. We did not run a controlled ablation isolating the exact percentage-point impact, so I will not put a hard number on it, but on top of activating the judge it is a real contributor here. Agpt-oss-120b/deepseek-v4-class judge is an equally good modern alternative.Q2. Frames (512 vs 768)
This is the config delta I would focus on, and the README's reproduce table is a little too terse about it. The paper's actual protocol (Experimental Setup, "Dense Frame Sampling" regime) is: 512 or 768 uniformly-sampled frames, and we report the better of the two. All three stages were trained at 512 frames (training-hyperparameters table), and the SFT checkpoint is highly frame-count-sensitive: the same SFT model scores only 12.5 under sparse 64-frame sampling vs 64.9 dense, a ~52-point swing.
You ran 768 only. Because the SFT checkpoint is train-aligned at 512 and so frame-sensitive, 768 may simply be the worse of the two for it. Please also run
max_frames=512and take the better of {512, 768}; that is what the reported number reflects. (For the RFT checkpoint, 512 is in fact our documented default, which is probably the source of any "512" guidance you have seen.) I will fix the FAQ to state the better-of-two protocol explicitly rather than a flat 768.Q3. Tool behavior on long videos
A few things are simultaneously true here, and I would rather be straight than defensive:
LongVT-7B-SFTis the cold-start checkpoint, stage one of a three-stage SFT → RL → RFT pipeline, and the weakest released stage (VideoMME dense: SFT 64.9 < RL 66.1 < RFT 67.0). The paper is explicit that disciplined tool use (when to crop, how long a window) is what RL (with an IoU reward that curbs span inflation) and RFT stabilize, and that SFT alone yields only modest long-video gains. At the SFT stage the tool fires but is not yet fully disciplined, so on VideoMME its net accuracy contribution is small and the 64.9 is carried mostly by dense-frame coverage; the tool-use payoff shows up after RL/RFT. So a smaller tool contribution at the SFT stage is consistent with the paper's own staging, where RL and RFT are what make tool use pay off. That does not excuse the SFT number (64.9 is published and you deserve to hit it), but the headline tool-use behavior lives in RFT.- Mechanism:
crop_video(video_path, start_time, end_time)takes absolute timestamps in seconds that the model itself predicts from its temporal grounding; there is no external timestamp or frame-time signal injected (the chat template emits only<|video_pad|>, no per-frame time tokens). The tool enforces onlystart < endand in-bounds timestamps; it does not impose a minimum window duration. On your "78% of crops start < 1 s" observation, I would be careful reading that as a failure mode, because early crop starts are common even in the healthy model: in the Questions about the evaluation. #19 RFT reference log (below), about 50% of crop start times are < 1 s (median 2 s, mean ~91 s), so a model frequently cropping from near the beginning is partly expected, not by itself a sign of breakage. Your SFT number being higher (78%) is plausible for a cold-start checkpoint with less refined windowing, but I would not over-attribute the accuracy gap to it. I no longer have theLongVT-7B-SFTJSONL (it was an intermediate checkpoint and I have since left the previous company), so I cannot reproduce your exact 78% on my side. - For a healthy reference, compare against the eval log I posted in Questions about the evaluation. #19 (gdrive): it is the RFT model on VideoMME w/sub + tool, it scores 66.96% (1,808 / 2,700), tools fire in 2,418 / 2,700 samples, and the
<tool_call>crop_video {...}traces show the timestamp-driven cropping the headline numbers come from. That makes a useful side-by-side: if your SFT tool calls look structurally different (for example crops that overshoot the video duration, or malformed<tool_call>JSON), grepping your JSONL forcrop_videoerrors against that reference will surface it. But as the metric decomposition above shows, the recoverable accuracy for the SFT checkpoint is in the judge and frame settings, not primarily in tool behavior.
Q4. Other settings (system prompt, judge voting, pixel budget)
- System prompt / tool framing: this is governed by the provided chat template (
examples/eval/tool_call_qwen2_5_vl.jinja) that vLLM applies at serving time, not by a hidden client-side knob. When the tool schema is passed, that template emits the# Toolssystem block (rendering yourexamples/video_tools/mcp_server.pytool definitions) and the hermes<tool_call>format, so as long as you serve with--chat-template tool_call_qwen2_5_vl.jinjaand--tool-call-parser hermes(both are inexamples/eval/run_eval.sh) your tool and system framing match ours. One detail worth knowing: thevideomme_w_subtitle_reward_tooltask delivers the question throughdoc_to_textand does not prepend the repo's longerSYSTEM_PROMPT(the explicit<think>...</think>then<answer>...</answer>instruction); but that is identical for our runs and yours, so it is not a source of divergence between us. There is no separateasync_openaiversion you need to match here. - Judge voting: there is no voting / self-consistency (no
n>1, no majority vote). The scorer makes a single judge call withmax_tokens=5(plus at most one fallback call on the raw answer string). - Pixel budget:
max_pixels=50176(= 224×224) is exactly right; it is the eval default and matches the training config (video_max_pixels: 50176), so train/test are aligned. No change needed.
What you already got right
vLLM 0.12.0 (yes, that is our recommended eval version, with the provided
tool_call_qwen2_5_vl.jinjachat template and--tool-call-parser hermes; other vLLM versions have incompatiblehermes_tool_parserbehavior, so keep 0.12.0), the task namevideomme_w_subtitle_reward_tool,is_qwen3_vl=False(correct, all three checkpoints are fine-tuned from Qwen2.5-VL-7B-Instruct),max_pixels=50176, andmax_new_tokens=49152are all correct. Your ~610 output tokens/sample means you are on average not truncating (truncation was the failure mode in #19, fixed by the49152override). One safeguard worth applying regardless: if you spot any individual responses cut off before the<answer>tag, raisemax_new_tokensfurther (for example 65536) and re-run, so every question is answered in full, since a truncated response scores 0 regardless of the judge.tp=4versus ourdp=8serving is a throughput choice and should not move the score.Suggested next steps (most actionable first)
- Turn the judge on (
USE_LLM_JUDGE=True, reachable endpoint). Youracc_score=0.3974is consistent with it being off; your own 53.67 with the 72B judge shows ~+14 pts from activating it alone. (An unreachable judge silently zeroes every non-exact answer.) - Use the
Qwen3-235B-A22Bjudge we used for the reported number (or agpt-oss-120b/deepseek-v4-class judge); the 72B → 235B swap closes part of the remaining 53.67 → 64.9 gap. - Run
max_frames=512as well and take the better of {512, 768}, matching the reported protocol; the SFT checkpoint is train-aligned at 512. - Mop up the format tail. Grep
acc_score=0samples for a correct letter in the reasoning but no clean<answer>(yourformat_scoreis 0.7256); with the judge on, the remainder is mostly genuine misses. - If your goal is the headline tool-use behavior, also evaluate
LongVT-7B-RFTand compare its trace against the Questions about the evaluation. #19 reference log, but the SFT 64.9 should be recoverable via steps 1–4.
Thanks again for the rigor; this is exactly the kind of report that improves the repo, and I will tighten the FAQ on the 512/768 "better-of-two" protocol and the judge/format notes. Happy to look at specific JSONL cases if you can share a handful.
- 32.78 (
- addedgood first issueGood for newcomersGood for newcomersquestionFurther information is requestedFurther information is requested
on Jun 24, 2026 Thanks, this clears up almost everything. Two quick clarifications so we can isolate the remaining gap:
- Judge was already on — my 53.67 is the
acc_scorepath withUSE_LLM_JUDGE=True+ a reachable 72B endpoint (the 39.74 was the judge-offresults.json), so I'm past step 1. - I ran both frame counts — 768→53.67, 512→54.74, best-of = 54.7 (the long-video bucket is what gains at 512). So frames is handled too.
I'll try more judge model(gpt-oss-120b / deepseek-v4-class) and re-score the deterministic-0 set without the >500-char gate.
Thank you again for your patient reply!
Reacted by Zuhao Yang- Judge was already on — my 53.67 is the
感谢你们出色的工作和详细的 FAQ。我按照 README/FAQ 操作,用官方的开源权重在VideoMME上进行测评,但始终比论文报告(表 2,“密集帧采样”,含字幕)中 LongVT-7B-SFT 在 VideoMME 上的 64.9% 低了约 11 %。我已经排除了 FAQ 中列出的常见问题,因此我将报告完整设置和已检查的内容,以防我遗漏了某个配置细节。
结果
按时长划分(评判准确率):短视频 65.4% / 中等视频 53.2% / 长视频 42.3%。短视频基本复现;全部差距都出现在中等视频,尤其是长视频中。这个结果在我约 7 次运行中都很稳定(得分都在 52–54% 之间)。
环境 / 确切配置
videomme_w_subtitle_reward_tool(已验证字幕确实被注入:2700 个样本中有 2232 个包含真实字幕文本;其余样本的数据集中没有.srt)--chat-template examples/eval/tool_call_qwen2_5_vl.jinja,--tool-call-parser hermes,--enable-auto-tool-choice,TP=4async_openai):fps=1, max_frames=768, max_pixels=50176, is_qwen3_vl=False, num_cpus=1, timeout=12000do_sample=False,temperature=0)max_new_tokens:49152(依据 FAQ Q2.5)accelerate --num_processes=4Qwen2.5-72B-Instruct,本地独立端口提供服务,USE_LLM_JUDGE=TrueFAQ 中已检查并确认不是原因的问题
videomme_w_subtitle_reward_tool(已确认提示中包含了字幕)。IS_QWEN3_VL=False。max_new_tokens=49152(只有 2.2% 的响应缺少<answer>标签或被截断)。--limit)。crop_video路径处理正常(使用干净的绝对路径;已配置前导斜杠的回退方案)。JSONDecodeError。问题
crop_video表现退化——78% 的裁剪起始时间 <1 秒(典型窗口为[0, ~32秒],而视频长度为 30–60 分钟),单次裁剪且无精细化,工具没有带来提升(无裁剪准确率 0.469 > 起始锚定裁剪准确率 0.408)。这是已发布 SFT 检查点的预期行为吗?还是说它应该能进行时间定位?另外,对于 Qwen2.5-VL(is_qwen3_vl=False)路径,全局帧是以无时间标记的独立image_url形式发送的,每帧的时间戳标记(<X seconds>)仅在 Qwen3-VL 路径中添加——那么 SFT 模型在评估时是否依赖任何时间戳/帧时间信号来确定裁剪时间?所使用的系统提示(来自
lmms_eval_tasks/videomme/utils.py):另外附上完整的推理 JSON 文件
20260204_234358_samples_videomme_reward_tool.json
20260622_102318_results.json
谢谢!