TurBLiMP is the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). This benchmark covers 16 core grammatical phenomena in Turkish, with 1,000 minimal pairs per phenomenon. Additionally, it incorporates experimental paradigms that examine model performance across different subordination strategies and word order variations.
TurBLiMP features 16 core phenomena:
- Anaphor Agreement - Reflexive pronoun agreement violations
- Argument Structure (Transitive) - Case marking errors with transitive verbs
- Argument Structure (Ditransitive) - Case marking errors with ditransitive verbs
- Binding - Principle B violations in binding theory
- Determiners - Obligatory use of the indefinite article
- Ellipsis - Backward gapping with non-parallel word orders
- Irregular Forms - Incorrect aorist allomorph usage
- Island Effects - Wh-adjunct extraction from complex NPs
- Nominalization - Incorrect nominalization suffix selection
- NPI Licensing - Negative polarity items in non-negative contexts
- Passives - Unlicensed use of by-phrases in impersonal passives
- Quantifiers - Quantifier usage with bare nouns
- Relative Clauses - Incorrect case marking in relative clauses
- Scrambling - Illicit postverbal scrambling from embedded clauses
- Subject Agreement - Person/number agreement violations
- Suspended Affixation - Improper tense suffix suspension
We also include 20 experimental paradigms targeting the ditransitive and transitive Argument Structure phenomena:
- Word Order - SOV, SVO, OSV, OVS, VSO, VOS
- Subordination - Finite, -DIK, -(y)IncA, -(y)ken
To validate our benchmark, we collected acceptability judgments from native speakers.
- Participants: 30 native Turkish speakers
- Rating scale: 7-point Likert (1: completely unacceptable – 7: completely acceptable)
- Stimuli:
- 216 sentences total
- Covers 16 linguistic phenomena and 20 experimental paradigms
- 6 sentences per category
- Design:
- Online survey on Qualtrics
- Two survey versions with flipped acceptability conditions
- Data included:
- Raw ratings for all phenomena
- Survey materials
To use the TurBLiMP benchmark:
- Clone the repository:
git clone https://github.com/yourusername/TurBLiMP.gitIf you use this benchmark in your work, please cite the associated paper.
@misc{başar2025turblimpturkishbenchmarklinguistic,
title={TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs},
author={Ezgi Başar and Francesca Padovani and Jaap Jumelet and Arianna Bisazza},
year={2025},
eprint={2506.13487},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.13487},
}
This work is licensed under a Creative Commons Attribution 4.0 International License.
