Skip to content

[Fix] Compare TydiQA answers case-insensitively - #2660

Open
MohammadHijjawi97 wants to merge 1 commit into
open-compass:mainfrom
MohammadHijjawi97:fix/tydiqa-case-insensitive-match
Open

MohammadHijjawi97 wants to merge 1 commit into
open-compass:mainfrom
MohammadHijjawi97:fix/tydiqa-case-insensitive-match

Conversation

@MohammadHijjawi97

Copy link
Copy Markdown

Motivation

TydiQAEvaluator.score lowercases each prediction before comparing it, but the gold answers are compared as-is, since general_postprocess does not lowercase. So a gold answer containing an uppercase letter can never reach exact match, and F1 drops every capitalized token, even when the model reproduces the reference verbatim:

TydiQAEvaluator().score(['Paris'], [['Paris']])
# {'exact_match': 0.0, 'f1': 0.0}

This affects all TyDiQA-GoldP languages with letter case (English, Russian, Finnish, Indonesian, Swahili, ...). The SQuAD v1.1 script this evaluator is adapted from lowercases both sides.

Modification

Lowercase the reference answers alongside the prediction in TydiQAEvaluator.score. Add tests/datasets/test_tydiqa.py, covering exact match / F1 with capitalized references and a wrong answer.

BC-breaking (Optional)

No API change. TyDiQA scores for cased languages will go up, since correct answers are no longer rejected because of case.

Use cases (Optional)

N/A

Checklist

Before PR:

  • Pre-commit or other linting tools are used to fix the potential lint issues. (flake8, isort, yapf and codespell on the changed files)
  • Bug fixes are fully covered by unit tests, the case that causes the bug should be added in the unit tests.
  • The modification is covered by complete unit tests. If not, please add more unit test to ensure the correctness.
  • The documentation has been modified accordingly, like docstring or example tutorials.

After PR:

  • If the modification has potential influence on downstream or other related projects, this PR should be tested with those projects.
  • CLA has been signed and all committers have signed the CLA in this PR.

TydiQAEvaluator lowercased the prediction but not the reference
answers, so any gold answer containing an uppercase letter could never
reach exact match and lost F1 on every capitalized token, even when the
model output matched the reference verbatim. Lowercase the references as
well, matching the SQuAD normalization the evaluator is based on.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant