Turkish Bilateral Investment Treaties Corpus (Dataset 1) is the machine-readable foundation of a three-dataset research corpus on Türkiye's international investment agreements. The current release contains 141 agreements signed from June 1962 through April 2024, with entry-into-force, termination and replacement status tracked through 31 December 2025.
Coverage & contents
The corpus covers 139 bilateral investment treaties, the 2005 Agreement on Promotion and Protection of Investments among ECO Member States, and the 2013 memorandum of understanding with the Islamic Development Bank. Every agreement has a stable identifier from TUR_BIT_001 to TUR_BIT_141.
- Full texts:
treaty_texts.zip contains 141 plain-text UTF-8 files, one per agreement, with a verified total of 432,055 whitespace-delimited words.
- Treaty metadata:
treaty_metadata.csv is a 141 × 22 treaty-level table covering identity, partner, dates, status, language, source, generation and protocol fields. The Parquet copy is cell-for-cell equivalent.
- Article extracts:
article_extracts.csv is a 141 × 71 wide table containing treaty identity and source fields, the preamble, paired title-and-text fields for up to 33 articles, and protocol material.
- FPS classification:
fps_classification.csv is a 141 × 7 treaty-level table recording the Pehlivan Full Protection and Security typology, formulation variant, article location and supporting wording.
- Documentation and code: the release includes a codebook, schema, reproducibility protocol, verification and rebuild scripts, and the canonical audit ledger
AUDIT_FINDINGS_2026-08-22.csv.
At the 31 December 2025 status cutoff, 103 agreements are in force, 23 are signed but not in force, 13 are terminated and 2 are replaced. The historical-generation distribution is 2 early, 55 liberalization, 29 EU-harmonization and 55 new-model agreements.
Methodology & data quality
Treaty texts were assembled primarily from officially promulgated sources, especially the Resmî Gazete, with official national and international repositories used for gap-filling and cross-checking. Text-native PDFs were extracted directly; scanned documents were processed with OCR and manually reviewed. The resulting texts are normalized to plain UTF-8 while preserving treaty structure and substantive wording.
Source fidelity is distinguished from transcription correction. A dataset transfer error was corrected only where the archived text diverged from the visible source PDF. The final audit made 10 such PDF-verified OCR or transfer fixes across seven treaties. By contrast, 55 unusual instances in 37 treaties were confirmed as actually printed in the source PDFs and were deliberately retained, even where the wording, spacing or glyph appears erroneous. These retained instances document the source rather than silently modernizing it.
The same audit repaired 76 truncated article headings and 24 duplicated title/body boundaries only in Dataset 1 article_extracts.csv and the corresponding Dataset 2 clause_extracts.csv; these were extraction-structure errors, not changes to treaty substance. Separately, 10 PDF-verified text repairs were synchronized from treaty_texts.zip through article_extracts.csv to clause_extracts.csv. Dataset 3 was then regenerated from the audited sources.
Quality controls verify unique and complete treaty identifiers, valid UTF-8 and ZIP integrity, source-filename mappings, metadata word and article counts, CSV/Parquet equivalence, traceability of extracted articles to the full texts, and synchronization of FPS fields across the three datasets. The deposited cross-dataset verification suite passes on the current release.
Key findings
The Pehlivan FPS Typology classifies every agreement by the relationship between Full Protection and Security, fair and equitable treatment, the minimum standard of treatment, applicable-law qualifications and comparator standards. The audited distribution is:
- Type A — standalone FPS: 8 agreements.
- Type B — FET plus FPS: 31 agreements.
- Type C — minimum standard of treatment plus FET and FPS: 40 agreements.
- Type D — domestic-law reference plus FPS: 2 agreements.
- Type E — international-law reference plus FET and FPS: 10 agreements.
- Type F — national-treatment or most-favoured-nation comparator plus FPS: 5 agreements.
- Type G — no FPS clause: 45 agreements.
Type G is the largest category, while Type C is the dominant new-model formulation. The corpus also shows a shift in clause placement from Article 2 in older agreements toward Articles 3 and 4 in newer agreements. Eleven agreements signed from 2015 onward expressly limit FPS to physical or police protection, either in the treatment article or in an attached interpretive paragraph.
Reproducibility & reuse
00_REPRODUCIBILITY_PROTOCOL.txt documents the complete workflow from source texts through metadata, article extraction, FPS classification, Dataset 2 clause coding and Dataset 3 network reconstruction. The code/ directory contains the corresponding verification, audit and rebuild scripts. Use treaty_id as the primary join key across all three datasets; use the metadata source field to map a treaty to its member in treaty_texts.zip.
The DataverseNO record is the maintained version of record: 10.18710/JX4WHH. An all-versions mirror is available through Zenodo at 10.5281/zenodo.20492363. Development takes place at github.com/ouzpehlivan/turkish-bit-corpus; the repository may run ahead of the preserved deposit.
Ethics & licence
The research involved no human participants and no data collected from individuals, so no informed consent, research-ethics committee approval or notification to Sikt was required. Personal data are limited to the names and official titles of public officials appearing in their official capacity in promulgated treaty texts; the analytical DS1 tables do not introduce private-person or special-category data.
The dataset and documentation are released under Creative Commons Attribution 4.0 International. The reproducibility software in code/ is released under the MIT licence documented in code/LICENSE_MIT.txt. Official treaty texts remain official legal acts of the contracting states; source and terms-of-use details are documented in 00_README.txt.
Related datasets & DOIs
Dataset 1 — Turkish BIT Treaty Corpus: this dataset, DataverseNO DOI 10.18710/JX4WHH.
Dataset 2 — Turkish BIT Clause-Level Annotation Dataset: 118 substantive variables coded per treaty, DataverseNO DOI 10.18710/FPNLKS.
Dataset 3 — Turkish BIT Treaty Diffusion and Genealogy Network: pairwise textual similarity, treaty families and genealogy edges derived from Datasets 1 and 2, DataverseNO DOI 10.18710/WA7HEO.
(2026-08-24)