Skip to content

Question about reproducing the Super-Resolution / SRIQA-Bench results in Table 2 #15

Description

@jiangsiqi466-debug

Hi, thank you for releasing VisualQuality-R1 and the detailed README.

I am trying to reproduce the OOD evaluation result of VisualQuality-R1 on the Super-Resolution / SRIQA-Bench column in Table 2. However, I observe a large gap between my reproduced result and the reported result.
In Table 2, VisualQuality-R1 reports approximately:

  • Super-Resolution SRCC: 0.756
  • Super-Resolution PLCC: 0.825

In my local evaluation, using the released VisualQuality-R1-7B checkpoint and the README non-thinking prompt, I obtain much lower results on my SRIQA-Bench label file.
My current setting is:

  • Model: released VisualQuality-R1-7B checkpoint
  • Prompt: the README non-thinking prompt, with rating between 1 and 5 and output in tags
  • Decoding: do_sample=False
  • Evaluation format: no-reference single-image scoring
  • Label format: ref_image restored_image mos
  • Number of samples: 1100
  • Image input: the second column restored image
  • Ground truth: the third column MOS score

An example line in my SRIQA-Bench label file is:
online50_Original.png online50_SeeSRx4.png 100.0

I parse online50_SeeSRx4.png as the evaluated image and 100.0 as the ground-truth score.
I have checked several possible causes:

  1. The SRIQA-Bench label is parsed as restored_image + MOS, not the reference image.
  2. The README prompt is indeed a 1–5 rating prompt, so I use the same prompt style.
  3. I tested larger image resolutions. When increasing max_side from 448 to 896, the result improves slightly, but it is still far below the reported Table 2 result.

My current results are approximately:

  • max_side=448: SRCC 0.555 / PLCC 0.546
  • max_side=896: SRCC 0.572 / PLCC 0.557

Therefore, I would like to ask about the exact evaluation protocol used for the Super-Resolution / SRIQA-Bench result in Table 2.
Could you please clarify the following questions?

  1. For the Super-Resolution result in Table 2, did you use the full SRIQA-Bench set? If so, how many samples were evaluated?
  2. Does the evaluation include the original images as test images? In my local file, there are 1100 samples, which seems to include 100 original images plus 1000 SR results. Or should only the 1000 SR images be evaluated?
  3. Is the Super-Resolution evaluation performed as no-reference single-image scoring, or does the model use both the reference image and the restored/SR image?
  4. Could you share the exact SRIQA-Bench label file format used in Table 2?
  5. Did you resize images during evaluation? If so, what image resolution, max_side, or max_pixels setting was used?
  6. For PLCC, did you compute Pearson correlation directly, or did you apply nonlinear logistic fitting before computing PLCC?
  7. Is the released VisualQuality-R1-7B checkpoint exactly the same checkpoint used to report the Table 2 Super-Resolution result, or was Table 2 evaluated with an internal checkpoint trained with additional private data?
  8. Would it be possible to provide the evaluation script or label file used for the SRIQA-Bench / Super-Resolution result in Table 2?

I understand that the final version of VisualQuality-R1 may include some private training data. I mainly want to make sure whether my SRIQA-Bench evaluation protocol is aligned with yours.

Thank you very much for your time and help!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions