Hi, thank you for releasing MBench.
We are evaluating the 227 released LingBot-World items with
subset = environment in
MBench-A/models/lingbot_world/samples.jsonl, using the public code at
commit c44dbd9.
Using the released videos, DA3-GIANT-1.1 over all frames, and the public
mbench eval pipeline, we obtain:
| Metric |
Table 2 |
Ours |
| Epipolar |
22.12 |
0.13 |
| Reprojection |
7.57 |
1.14 |
| Lighting |
40.06 |
89.62 |
| Style |
85.87 |
87.06 |
All samples pass mbench validate. The rollout videos are byte-identical
to the public release.
Could you please share the exact DA3 artifact generation setup used for
Table 2 (checkpoint, preprocessing, frame sampling, pose convention), or
release the LingBot-World results.npz and corresponding input_images?
Also, could you confirm whether Table 2 used the current public evaluator
and the same released environment subset?
Thank you!
Hi, thank you for releasing MBench.
We are evaluating the 227 released LingBot-World items with
subset = environmentinMBench-A/models/lingbot_world/samples.jsonl, using the public code atcommit
c44dbd9.Using the released videos, DA3-GIANT-1.1 over all frames, and the public
mbench evalpipeline, we obtain:All samples pass
mbench validate. The rollout videos are byte-identicalto the public release.
Could you please share the exact DA3 artifact generation setup used for
Table 2 (checkpoint, preprocessing, frame sampling, pose convention), or
release the LingBot-World
results.npzand correspondinginput_images?Also, could you confirm whether Table 2 used the current public evaluator
and the same released
environmentsubset?Thank you!