{"id":188565,"identifier":"5KCE4U","persistentUrl":"https://doi.org/10.18710/5KCE4U","protocol":"doi","authority":"10.18710","separator":"/","publisher":"DataverseNO","publicationDate":"2023-10-24","storageIdentifier":"S3://10.18710/5KCE4U","datasetType":"dataset","datasetVersion":{"id":4779,"datasetId":188565,"datasetPersistentId":"doi:10.18710/5KCE4U","storageIdentifier":"S3://10.18710/5KCE4U","versionNumber":1,"versionMinorNumber":1,"versionState":"RELEASED","latestVersionPublishingState":"RELEASED","deaccessionLink":"","lastUpdateTime":"2025-07-17T06:55:34Z","releaseTime":"2025-07-17T06:55:34Z","createTime":"2025-07-16T19:54:33Z","publicationDate":"2023-10-24","citationDate":"2023-10-24","termsOfUse":"<p>With the exception of the tabular file data_jenset_mcgillivray_downsampling.tsv, the dataset “Background data (adapted from Jenset & McGillivray 2017) for: Down-sampling from hierarchically structured corpus data” has been marked as dedicated to the public domain, as described here: <a href=\"https://creativecommons.org/publicdomain/zero/1.0/\">https://creativecommons.org/publicdomain/zero/1.0/</a>.</p>\n<p>Our <a href=\"https://dataverse.org/best-practices/dataverse-community-norms\">Community Norms</a> as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.</p>\n<p>The tabular file data_jenset_mcgillivray_downsampling.tsv constitutes an adaptation of a dataset published by Gard Jenset on GitHub in 2018 (\"Jenset 2018\"), available at <a href=\"https://github.com/gjenset/quanthistbook/tree/master/eme_v3sng_study\n\">https://github.com/gjenset/quanthistbook/tree/master/eme_v3sng_study</a>. Material has been adapted from Jenset 2018, as documented in the ReadMe file included in the present dataset, under a MIT License. Contributions made by the author of the present dataset to the adaptation are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, as described here: <a href=\"https://creativecommons.org/licenses/by/4.0/\">https://creativecommons.org/licenses/by/4.0/</a>. Reusers of the adapted work must comply with both this CC license and the original MIT License:\n<blockquote>\n<p>Copyright (c) 2018 Gard Jenset</p>\n<p>Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:</p>\n<p>The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.</p>\n<p>THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>\n</blockquote>\n</p>\n<p>Data included in Jenset 2018 were in turn derived from The Penn-Helsinki Parsed Corpus of Early Modern English (PPCEME), available here: <a href=\"https://www.ling.upenn.edu/ppche/ppche-release-2016/PPCEME-RELEASE-3\">https://www.ling.upenn.edu/ppche/ppche-release-2016/PPCEME-RELEASE-3</a>. Note that the MIT License under which Jenset 2018 is made available cannot be considered to apply to text fragments extracted from PPCEME.","fileAccessRequest":true,"metadataBlocks":{"citation":{"displayName":"Citation Metadata","name":"citation","fields":[{"typeName":"title","multiple":false,"typeClass":"primitive","value":"Background data (adapted from Jenset & McGillivray 2017) for: Down-sampling from hierarchically structured corpus data"},{"typeName":"author","multiple":true,"typeClass":"compound","value":[{"authorName":{"typeName":"authorName","multiple":false,"typeClass":"primitive","value":"Sönning, Lukas"},"authorAffiliation":{"typeName":"authorAffiliation","multiple":false,"typeClass":"primitive","value":"University of Bamberg"},"authorIdentifierScheme":{"typeName":"authorIdentifierScheme","multiple":false,"typeClass":"controlledVocabulary","value":"ORCID"},"authorIdentifier":{"typeName":"authorIdentifier","multiple":false,"typeClass":"primitive","value":"0000-0002-2705-395X"}}]},{"typeName":"datasetContact","multiple":true,"typeClass":"compound","value":[{"datasetContactName":{"typeName":"datasetContactName","multiple":false,"typeClass":"primitive","value":"Sönning, Lukas"},"datasetContactAffiliation":{"typeName":"datasetContactAffiliation","multiple":false,"typeClass":"primitive","value":"University of Bamberg"},"datasetContactEmail":{"typeName":"datasetContactEmail","multiple":false,"typeClass":"primitive","value":"lukas.soenning@uni-bamberg.de"}}]},{"typeName":"dsDescription","multiple":true,"typeClass":"compound","value":[{"dsDescriptionValue":{"typeName":"dsDescriptionValue","multiple":false,"typeClass":"primitive","value":"<p><strong>Dataset description</strong></p>\n<p>This dataset, which is adapted from Jenset and McGillivray (2017), contains tabular files documenting the alternating usage of -(e)th and -(e)s to mark third-person verb inflection in Early Modern English. The data provided by Jenset and McGillivray (2017) are drawn from the PPCEME corpus (Kroch et al. 2004) and cover the period from 1500 to 1700. In total, 13,757 third-person singular tokens (excluding the verb BE) were annotated by these authors for a range of variables. For the purposes of the present methodological study, this dataset was reduced to a subset of 11,645 tokens, and the coding of variables was in some parts revised, completed, or modified. The dataset includes information about the Author and Verb Lemma, as well as a number of predictor variables, including Genre, Year, Frequency (of the verb lemma in the third-person singular), Phonological Context (stem-final sound), and the Gender of the author.</p>"},"dsDescriptionDate":{"typeName":"dsDescriptionDate","multiple":false,"typeClass":"primitive","value":"2023-07-20"}},{"dsDescriptionValue":{"typeName":"dsDescriptionValue","multiple":false,"typeClass":"primitive","value":"<p><strong>Abstract for related publication</strong></p>\n<p>Resource constraints often force researchers to down-size the list of tokens returned by a corpus query. This paper sketches a methodology for down-sampling and offers a survey of current practices. We build on earlier work and extend the evaluation of down-sampling designs to settings where tokens are clustered by text file and lexeme. Our case study deals with third-person present-tense verb inflection in Early Modern English and focuses on five predictors: Year, Gender, Genre, Frequency, and Phonological Context. We evaluate two strategies for selecting 2,000 (out of 11,645) tokens: simple down-sampling, where each hit has the same selection probability; and structured down-sampling, where this probability is inversely proportional to the author- and verb-specific token count. We form 500 sub-samples using each scheme and compare regression results to a reference model fit to the full set of cases. We observe that structured down-sampling shows better performance on several evaluation criteria.</p>"},"dsDescriptionDate":{"typeName":"dsDescriptionDate","multiple":false,"typeClass":"primitive","value":"2023-10-23"}}]},{"typeName":"subject","multiple":true,"typeClass":"controlledVocabulary","value":["Arts and Humanities"]},{"typeName":"keyword","multiple":true,"typeClass":"compound","value":[{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"Early Modern English"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"verb inflection"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"language change"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"lexical diffusion"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"third person singular"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"methodology"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"down-sampling"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"corpus linguistics"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"PPCEME"}},{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"Penn-Helsinki Parsed Corpus of Early Modern English"}}]},{"typeName":"publication","multiple":true,"typeClass":"compound","value":[{"publicationCitation":{"typeName":"publicationCitation","multiple":false,"typeClass":"primitive","value":"Sönning, Lukas. 2024. Down-sampling from hierarchically structured corpus data. International Journal of Corpus Linguistics 29(4). 507–533."},"publicationIDType":{"typeName":"publicationIDType","multiple":false,"typeClass":"controlledVocabulary","value":"doi"},"publicationIDNumber":{"typeName":"publicationIDNumber","multiple":false,"typeClass":"primitive","value":"10.1075/ijcl.23079.son"},"publicationURL":{"typeName":"publicationURL","multiple":false,"typeClass":"primitive","value":"https://doi.org/10.1075/ijcl.23079.son"}}]},{"typeName":"language","multiple":true,"typeClass":"controlledVocabulary","value":["English"]},{"typeName":"producer","multiple":true,"typeClass":"compound","value":[{"producerName":{"typeName":"producerName","multiple":false,"typeClass":"primitive","value":"Alan Turing Institute, University of Cambridge"},"producerURL":{"typeName":"producerURL","multiple":false,"typeClass":"primitive","value":"https://www.c2d3.cam.ac.uk/research/alan-turing-institute"}},{"producerName":{"typeName":"producerName","multiple":false,"typeClass":"primitive","value":"University of Bamberg"},"producerURL":{"typeName":"producerURL","multiple":false,"typeClass":"primitive","value":"https://www.uni-bamberg.de/eng-ling/"}}]},{"typeName":"distributor","multiple":true,"typeClass":"compound","value":[{"distributorName":{"typeName":"distributorName","multiple":false,"typeClass":"primitive","value":"The Tromsø Repository of Language and Linguistics (TROLLing)"},"distributorAbbreviation":{"typeName":"distributorAbbreviation","multiple":false,"typeClass":"primitive","value":"TROLLing"},"distributorURL":{"typeName":"distributorURL","multiple":false,"typeClass":"primitive","value":"https://trolling.uit.no/"}}]},{"typeName":"depositor","multiple":false,"typeClass":"primitive","value":"Sönning, Lukas"},{"typeName":"dateOfDeposit","multiple":false,"typeClass":"primitive","value":"2023-07-20"},{"typeName":"timePeriodCovered","multiple":true,"typeClass":"compound","value":[{"timePeriodCoveredStart":{"typeName":"timePeriodCoveredStart","multiple":false,"typeClass":"primitive","value":"1500-01-01"},"timePeriodCoveredEnd":{"typeName":"timePeriodCoveredEnd","multiple":false,"typeClass":"primitive","value":"1707-12-31"}}]},{"typeName":"dateOfCollection","multiple":true,"typeClass":"compound","value":[{"dateOfCollectionStart":{"typeName":"dateOfCollectionStart","multiple":false,"typeClass":"primitive","value":"2022-11-15"},"dateOfCollectionEnd":{"typeName":"dateOfCollectionEnd","multiple":false,"typeClass":"primitive","value":"2023-06-15"}}]},{"typeName":"kindOfData","multiple":true,"typeClass":"primitive","value":["observational data","textual linguistic data","corpus data"]},{"typeName":"software","multiple":true,"typeClass":"compound","value":[{"softwareName":{"typeName":"softwareName","multiple":false,"typeClass":"primitive","value":"R"},"softwareVersion":{"typeName":"softwareVersion","multiple":false,"typeClass":"primitive","value":"4.2.1"}},{"softwareName":{"typeName":"softwareName","multiple":false,"typeClass":"primitive","value":"RStudio"},"softwareVersion":{"typeName":"softwareVersion","multiple":false,"typeClass":"primitive","value":"2023.06.2"}}]},{"typeName":"dataSources","multiple":true,"typeClass":"primitive","value":["<p>Data in the tabular file data_jenset_mcgillivray_downsampling.tsv has been adapted from a dataset published by Gard Jenset in 2018 (\"Jenset 2018\") on GitHub at <a href=\"https://github.com/gjenset/quanthistbook/tree/master/eme_v3sng_study\">https://github.com/gjenset/quanthistbook/tree/master/eme_v3sng_study.</a></p>\n<p>Jenset 2018 contains supporting data for Gard B. Jenset and Barbara McGillivray, 'A new methodology for quantitative historical linguistics', <i>Quantitative Historical Linguistics: A Corpus Framework</i>, Oxford Studies in Diachronic and Historical Linguistics (Oxford, 2017), <a href=\"https://doi.org/10.1093/oso/9780198718178.003.0007\">https://doi.org/10.1093/oso/9780198718178.003.0007</a>.</p>\n<p>Data from Jenset 2018 is reused here (as described in the ReadMe file) under a MIT License:</p>\n<blockquote><p>Copyright (c) 2018 Gard Jenset</p>\n<p>Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:</p>\n<p>The above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.</p>\n<p>THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.</p>\n</blockquote>\n<p>Data included in Jenset 2018 were in turn derived from: Anthony Kroch, Beatrice Santorini, and Lauren Delfs. 2004. The Penn-Helsinki Parsed Corpus of Early Modern English (PPCEME). Department of Linguistics, University of Pennsylvania. <a href=\"https://www.ling.upenn.edu/ppche/ppche-release-2016/PPCEME-RELEASE-3\">https://www.ling.upenn.edu/ppche/ppche-release-2016/PPCEME-RELEASE-3</a>.</p>\n<p>PPCEME is currently distributed by the Linguistic Data Consortium as part of the Penn Parsed Corpora of Historical English under a user license agreement. This agreement permits the User to \"include limited excerpts from the Data in articles, reports and other documents describing the results of User’s non-commercial projects related to linguistic education, research and technology development\". The text fragments extracted from PPCEME by Jenset and McGillivray and incorporated into the present dataset only represent limited excerpts of the kind that may be shared under limitations and exceptions to copyright, such as Fair Use or Fair Dealing.</p>"]}]},"geospatial":{"displayName":"Geospatial Metadata","name":"geospatial","fields":[{"typeName":"geographicCoverage","multiple":true,"typeClass":"compound","value":[{"country":{"typeName":"country","multiple":false,"typeClass":"controlledVocabulary","value":"United Kingdom"}}]}]}},"files":[{"description":"Text file describing the dataset","label":"00_ReadMe_downsampling.txt","restricted":false,"version":1,"datasetVersionId":4779,"categories":["Documentation"],"dataFile":{"id":190377,"persistentId":"doi:10.18710/5KCE4U/GXUMC0","pidURL":"https://doi.org/10.18710/5KCE4U/GXUMC0","filename":"00_ReadMe_downsampling.txt","contentType":"text/plain","friendlyType":"Plain Text","filesize":12381,"description":"Text file describing the dataset","categories":["Documentation"],"storageIdentifier":"S3://uit-dataverseno-prod01:18b5d07ef40-5d3a3214bfb6","rootDataFileId":-1,"md5":"d2b604408fd1f1ffb99c28cfc66c3760","checksum":{"type":"MD5","value":"d2b604408fd1f1ffb99c28cfc66c3760"},"tabularData":false,"creationDate":"2023-10-23","publicationDate":"2023-10-24","fileAccessRequest":true}},{"description":"Tab-delimited data table containing the 11,645 annotated verb tokens","label":"data_jenset_mcgillivray_downsampling.tsv","restricted":false,"version":1,"datasetVersionId":4779,"categories":["Data"],"dataFile":{"id":190376,"persistentId":"doi:10.18710/5KCE4U/LJVY2I","pidURL":"https://doi.org/10.18710/5KCE4U/LJVY2I","filename":"data_jenset_mcgillivray_downsampling.tsv","contentType":"text/tsv","friendlyType":"Tab-Separated Values","filesize":2120816,"description":"Tab-delimited data table containing the 11,645 annotated verb tokens","categories":["Data"],"storageIdentifier":"S3://uit-dataverseno-prod01:18b5ce624c9-96de46d0c5fb","rootDataFileId":-1,"md5":"ceca38119ea898d2bbb23c5dd5a32b41","checksum":{"type":"MD5","value":"ceca38119ea898d2bbb23c5dd5a32b41"},"tabularData":false,"creationDate":"2023-10-23","publicationDate":"2023-10-24","fileAccessRequest":true}},{"description":"R script documenting the data preparation steps","label":"script_data_preparation.qmd","restricted":false,"version":1,"datasetVersionId":4779,"categories":["Code"],"dataFile":{"id":189693,"persistentId":"doi:10.18710/5KCE4U/YGMFU7","pidURL":"https://doi.org/10.18710/5KCE4U/YGMFU7","filename":"script_data_preparation.qmd","contentType":"application/octet-stream","friendlyType":"Unknown","filesize":13462,"description":"R script documenting the data preparation steps","categories":["Code"],"storageIdentifier":"S3://uit-dataverseno-prod01:18a97d8b74b-ea377c32bf56","rootDataFileId":-1,"md5":"78c7f9528d13d6bb2bddc354e6e62a19","checksum":{"type":"MD5","value":"78c7f9528d13d6bb2bddc354e6e62a19"},"tabularData":false,"creationDate":"2023-09-15","publicationDate":"2023-10-24","fileAccessRequest":true}}],"citation":"Sönning, Lukas, 2023, \"Background data (adapted from Jenset & McGillivray 2017) for: Down-sampling from hierarchically structured corpus data\", https://doi.org/10.18710/5KCE4U, DataverseNO, V1"}}