|
Description
| Turkish BIT Clause-Level Annotation Dataset (Dataset 2). This dataset provides a treaty-level, clause-aware coding of all 141 international investment agreements in the Republic of Türkiye corpus. The agreements were signed between June 1962 and April 2024, and treaty-status metadata are tracked through 31 December 2025. Dataset 2 is the analytical companion to Dataset 1: it converts the treaty texts into 118 documented substantive variables that can be compared across partners, periods and treaty generations. The maintained DataverseNO record is 10.18710/FPNLKS.
Coverage and contents
The main table, treaty_annotations.csv, has 141 rows and 122 columns: 118 substantive variables, treaty_id, and three provenance fields. Its Apache Parquet counterpart contains the same 141 × 122 cells in the same column order. The 118 variables comprise 101 parameters implementing the UNCTAD IIA Mapping Project methodology and 17 custom additions designed for Turkish treaty practice, including the Pehlivan Full Protection and Security typology.
treaty_annotations.csv and treaty_annotations.parquet: one fully coded row per treaty, joined to Datasets 1 and 3 by treaty_id.
clause_extracts.csv: 1,890 article-level analytical segments, each carrying the treaty, article, primary clause type, article title and source-faithful clause text.
coding_log.csv: one provenance and verification record per treaty.
ds2_change_log.csv: 1,283 row-level change records documenting the variable, previous value, revised value and reason.
CODEBOOK.txt and VARIABLES.txt: file, field, option, conditional-coding and validation documentation.
- Figures and explorer: five analytical figures, a static dashboard and an offline explorer covering all 117 map-compatible substantive variables. The explorer intentionally excludes only
3.13_fps_text, because a long verbatim clause is not a categorical mapping variable; the field remains fully available in the main table.
Coding methodology and data quality
Coding is context-aware and source-linked. Each value was derived from the cleaned treaty text and article extracts in Dataset 1, using the relevant operative provision rather than isolated keyword hits. Ambiguous and multi-function articles were adjudicated against the full article and treaty. Conditional fields distinguish None, where the mapped option is absent, from Not applicable, where the parent provision is absent.
The 22 August 2026 full-corpus audit re-read every treaty population-wide. It distinguished operative fair and equitable treatment from aspirational preambular wording; required an additional arbitral forum to be an actual submission option rather than merely an appointing authority; required an express public or open-hearing rule for arbitral-hearing transparency; and separated initial duration, renewal, unilateral-termination notice and post-termination survival. That audit produced 826 annotation-cell corrections across 67 variables. It also corrected 14 primary clause_type assignments in clause_extracts.csv. Earlier review rounds remain preserved in the 1,283-row change log rather than being overwritten. A follow-up review on 25 August 2026 closed the eighteen codings the audit had left flagged as unresolved, accepting five and confirming twelve, and re-derived 1.05_preamble_objective across all 141 preambles under an explicit decision rule. The rules adopted are documented in VARIABLES.txt under Operational decision rules.
- All 141 annotation rows are marked
OKP_verified and high confidence in both the main table and coding log.
- All 1,890 clause extracts are checked against their Dataset 1 source articles, and the controlled
clause_type distribution is recomputed during verification.
- The CSV and Parquet annotation tables have identical dimensions, columns and cell values.
- Dataset 1 and Dataset 2 treaty identifiers, partner metadata and FPS classifications are checked across files; dependent Dataset 3 outputs are regenerated rather than edited by hand.
Key findings
- Fair and equitable treatment: 53 treaties contain qualified FET, 46 contain unqualified FET and 42 contain no operative FET obligation. In 35 of the 42 None records, wording previously treated as FET appeared only in an aspirational preamble rather than an operative treatment article.
- National treatment: 124 treaties provide post-establishment national treatment, 15 extend it to the pre-establishment phase and 2 contain no national-treatment clause.
- ISDS and ICSID: 139 treaties provide investor–State dispute settlement and 133 offer ICSID Convention arbitration. Germany 1962 and the Islamic Development Bank memorandum are the two treaties without ISDS.
- Scope of ISDS claims: 94 treaties cover any dispute relating to an investment, 13 list specific bases of claim beyond the treaty, 32 cover treaty claims only and 2 are Not applicable.
- Additional arbitral forums: 64 treaties offer another institution or set of rules beyond domestic courts, ICSID Convention arbitration and UNCITRAL arbitration; 75 do not and 2 are Not applicable.
- Relationship between forums: 64 treaties use a fork-in-the-road rule, 38 preserve arbitration after domestic proceedings under specified conditions, 26 contain no reference, 6 require local remedies first, 5 use a no-U-turn or waiver rule and 2 are Not applicable.
- Public hearings: only the Lithuania 2018 treaty,
TUR_BIT_123, expressly requires public or open arbitral hearings. The final distribution is 1 Yes, 138 No and 2 Not applicable.
- Umbrella clauses and denial of benefits: 18 treaties contain an investment-directed umbrella clause, all with broad any-obligation wording, and 44 contain a denial-of-benefits clause.
- ESG-related language: 60 treaties contain at least one provision in the documented ESG composite, which covers sustainable-development, social or environmental preambular language; selected operative health, environment, labour, corporate-social-responsibility and not-lowering clauses; and general health or environmental exceptions. This is a descriptive composite and does not itself establish an enforceable ESG obligation.
- Duration and termination: initial terms are ten years in 128 treaties, fifteen years in 11, indefinite in 1 and other in 1. Renewal is indefinite in 129 treaties, ten years in 5, two years in 2, five years in 2, fifteen years in 1, other in 1 and none in 1. All 141 treaties specify unilateral-termination modalities; notice is one year in 132, six months in 7 and another period in 2. Amendment modalities appear in 124 treaties. Survival periods are ten years in 120 treaties, fifteen years in 14, five years in 5, twenty years in 1 and other in 1.
- Full protection and security: the Pehlivan FPS distribution is A=8, B=31, C=40, D=2, E=10, F=5 and G=45, identical to Dataset 1.
Reproducibility and reuse
The Python reproduction and verification toolkit is deposited with Dataset 1 in its code folder. It validates row and identifier coverage, CSV–Parquet equality, conditional values, FPS alignment, clause-source containment and cross-dataset dependencies; it also rebuilds the Dataset 2 figures and explorer and regenerates Dataset 3 from Dataset 1 texts and Dataset 2 annotations. Users should join files by treaty_id, consult VARIABLES.txt before aggregating categorical values, and use coding_log.csv and ds2_change_log.csv when provenance matters.
Development takes place at github.com/ouzpehlivan/turkish-bit-corpus. The repository is the working copy and may run ahead of the maintained DataverseNO deposits; archival citation should therefore use the dataset DOI and deposited version.
Ethics and licence
The dataset involves no human participants and no data collected from individuals for research purposes. No informed consent, research ethics committee approval or notification to Sikt was required. The only personal data are the names and official titles of state representatives reproduced in their public official capacity from promulgated treaty texts.
The dataset is released under Creative Commons Attribution 4.0 International. The reproduction code deposited with Dataset 1 is released under the MIT Licence. Bundled third-party visualisation libraries retain their own notices in THIRD_PARTY_LICENSES.txt.
Related datasets and identifiers
- Dataset 1: Turkish BIT Treaty Corpus. Cleaned treaty texts, metadata, article extracts and FPS classification. DataverseNO DOI: 10.18710/JX4WHH.
- Dataset 2: Turkish BIT Clause-Level Annotation Dataset. This dataset. DataverseNO DOI: 10.18710/FPNLKS.
- Dataset 3: Turkish BIT Treaty Diffusion and Genealogy Network. Pairwise textual and legal similarity, treaty families and genealogy edges. DataverseNO DOI: 10.18710/WA7HEO.
(2026-08-24) |