Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

TurBLiMP

TurBLiMP is the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). This benchmark covers 16 core grammatical phenomena in Turkish, with 1,000 minimal pairs per phenomenon. Additionally, it incorporates experimental paradigms that examine model performance across different subordination strategies and word order variations.

Table of Contents

📁 Linguistic Phenomena

TurBLiMP features 16 core phenomena:

  1. Anaphor Agreement - Reflexive pronoun agreement violations
  2. Argument Structure (Transitive) - Case marking errors with transitive verbs
  3. Argument Structure (Ditransitive) - Case marking errors with ditransitive verbs
  4. Binding - Principle B violations in binding theory
  5. Determiners - Obligatory use of the indefinite article
  6. Ellipsis - Backward gapping with non-parallel word orders
  7. Irregular Forms - Incorrect aorist allomorph usage
  8. Island Effects - Wh-adjunct extraction from complex NPs
  9. Nominalization - Incorrect nominalization suffix selection
  10. NPI Licensing - Negative polarity items in non-negative contexts
  11. Passives - Unlicensed use of by-phrases in impersonal passives
  12. Quantifiers - Quantifier usage with bare nouns
  13. Relative Clauses - Incorrect case marking in relative clauses
  14. Scrambling - Illicit postverbal scrambling from embedded clauses
  15. Subject Agreement - Person/number agreement violations
  16. Suspended Affixation - Improper tense suffix suspension

🔎 Experimental Paradigms

We also include 20 experimental paradigms targeting the ditransitive and transitive Argument Structure phenomena:

  1. Word Order - SOV, SVO, OSV, OVS, VSO, VOS
  2. Subordination - Finite, -DIK, -(y)IncA, -(y)ken

🙋 Human Judgments

To validate our benchmark, we collected acceptability judgments from native speakers.

  • Participants: 30 native Turkish speakers
  • Rating scale: 7-point Likert (1: completely unacceptable – 7: completely acceptable)
  • Stimuli:
    • 216 sentences total
    • Covers 16 linguistic phenomena and 20 experimental paradigms
    • 6 sentences per category
  • Design:
    • Online survey on Qualtrics
    • Two survey versions with flipped acceptability conditions
  • Data included:
    • Raw ratings for all phenomena
    • Survey materials

💻 Usage

To use the TurBLiMP benchmark:

  • Clone the repository:
git clone https://github.com/yourusername/TurBLiMP.git

🔗 Attribution

If you use this benchmark in your work, please cite the associated paper.

@misc{başar2025turblimpturkishbenchmarklinguistic,
      title={TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs}, 
      author={Ezgi Başar and Francesca Padovani and Jaap Jumelet and Arianna Bisazza},
      year={2025},
      eprint={2506.13487},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2506.13487}, 
}

🔓 License

CC BY 4.0

This work is licensed under a Creative Commons Attribution 4.0 International License.

CC BY 4.0

About

A Turkish Benchmark of Linguistic Minimal Pairs

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages