Skip to content

Cannot reproduce Table 2 LingBot-World scores #9

Description

@theREALevan

Hi, thank you for releasing MBench.

We are evaluating the 227 released LingBot-World items with
subset = environment in
MBench-A/models/lingbot_world/samples.jsonl, using the public code at
commit c44dbd9.

Using the released videos, DA3-GIANT-1.1 over all frames, and the public
mbench eval pipeline, we obtain:

Metric Table 2 Ours
Epipolar 22.12 0.13
Reprojection 7.57 1.14
Lighting 40.06 89.62
Style 85.87 87.06

All samples pass mbench validate. The rollout videos are byte-identical
to the public release.

Could you please share the exact DA3 artifact generation setup used for
Table 2 (checkpoint, preprocessing, frame sampling, pose convention), or
release the LingBot-World results.npz and corresponding input_images?

Also, could you confirm whether Table 2 used the current public evaluator
and the same released environment subset?

Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions