<resource xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://datacite.org/schema/kernel-4" xsi:schemaLocation="http://datacite.org/schema/kernel-4 http://schema.datacite.org/meta/kernel-4.1/metadata.xsd"><identifier identifierType="DOI">10.18710/5KCE4U</identifier><creators><creator><creatorName nameType="Personal">Sönning, Lukas</creatorName><givenName>Lukas</givenName><familyName>Sönning</familyName><nameIdentifier nameIdentifierScheme="ORCID">0000-0002-2705-395X</nameIdentifier><affiliation>University of Bamberg</affiliation></creator></creators><titles><title>Background data (adapted from Jenset &amp; McGillivray 2017) for: Down-sampling from hierarchically structured corpus data</title></titles><publisher>DataverseNO</publisher><publicationYear>2023</publicationYear><subjects><subject>Arts and Humanities</subject><subject>Early Modern English</subject><subject>verb inflection</subject><subject>language change</subject><subject>lexical diffusion</subject><subject>third person singular</subject><subject>methodology</subject><subject>down-sampling</subject><subject>corpus linguistics</subject><subject>PPCEME</subject><subject>Penn-Helsinki Parsed Corpus of Early Modern English</subject></subjects><contributors><contributor contributorType="ContactPerson"><contributorName nameType="Personal">Sönning, Lukas</contributorName><givenName>Lukas</givenName><familyName>Sönning</familyName><affiliation>University of Bamberg</affiliation></contributor><contributor contributorType="Producer"><contributorName nameType="Organizational">Alan Turing Institute, University of Cambridge</contributorName></contributor><contributor contributorType="Producer"><contributorName nameType="Organizational">University of Bamberg</contributorName></contributor><contributor contributorType="Distributor"><contributorName nameType="Personal">The Tromsø Repository of Language and Linguistics (TROLLing)</contributorName><givenName>The</givenName><familyName>Tromsø Repository of Language and Linguistics (TROLLing)</familyName></contributor></contributors><dates><date dateType="Submitted">2023-07-20</date><date dateType="Updated">2025-07-17</date><date dateType="Collected">2022-11-15/2023-06-15</date></dates><resourceType resourceTypeGeneral="Dataset">observational data</resourceType><relatedIdentifiers><relatedIdentifier relationType="IsCitedBy" relatedIdentifierType="DOI">10.1075/ijcl.23079.son</relatedIdentifier></relatedIdentifiers><sizes><size>12381</size><size>2120816</size><size>13462</size></sizes><formats><format>text/plain</format><format>text/tsv</format><format>application/octet-stream</format></formats><version>1.1</version><rightsList><rights rightsURI="info:eu-repo/semantics/openAccess"/><rights/></rightsList><descriptions><description descriptionType="Abstract">&lt;p>&lt;strong>Dataset description&lt;/strong>&lt;/p>
&lt;p>This dataset, which is adapted from Jenset and McGillivray (2017), contains tabular files documenting the alternating usage of -(e)th and -(e)s to mark third-person verb inflection in Early Modern English. The data provided by Jenset and McGillivray (2017) are drawn from the PPCEME corpus (Kroch et al. 2004) and cover the period from 1500 to 1700. In total, 13,757 third-person singular tokens (excluding the verb BE) were annotated by these authors for a range of variables. For the purposes of the present methodological study, this dataset was reduced to a subset of 11,645 tokens, and the coding of variables was in some parts revised, completed, or modified. The dataset includes information about the Author and Verb Lemma, as well as a number of predictor variables, including Genre, Year, Frequency (of the verb lemma in the third-person singular), Phonological Context (stem-final sound), and the Gender of the author.&lt;/p></description><description descriptionType="Abstract">&lt;p>&lt;strong>Abstract for related publication&lt;/strong>&lt;/p>
&lt;p>Resource constraints often force researchers to down-size the list of tokens returned by a corpus query. This paper sketches a methodology for down-sampling and offers a survey of current practices. We build on earlier work and extend the evaluation of down-sampling designs to settings where tokens are clustered by text file and lexeme. Our case study deals with third-person present-tense verb inflection in Early Modern English and focuses on five predictors: Year, Gender, Genre, Frequency, and Phonological Context. We evaluate two strategies for selecting 2,000 (out of 11,645) tokens: simple down-sampling, where each hit has the same selection probability; and structured down-sampling, where this probability is inversely proportional to the author- and verb-specific token count. We form 500 sub-samples using each scheme and compare regression results to a reference model fit to the full set of cases. We observe that structured down-sampling shows better performance on several evaluation criteria.&lt;/p></description><description descriptionType="TechnicalInfo">R, 4.2.1</description><description descriptionType="TechnicalInfo">RStudio, 2023.06.2</description></descriptions><geoLocations/></resource>