<?xml version='1.0' encoding='UTF-8'?><codeBook xmlns="ddi:codebook:2_5" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="ddi:codebook:2_5 https://ddialliance.org/Specification/DDI-Codebook/2.5/XMLSchema/codebook.xsd" version="2.5"><docDscr><citation><titlStmt><titl>Background data for: Advancing our understanding of dispersion measures in corpus research</titl><IDNo agency="DOI">doi:10.18710/FVHTFM</IDNo></titlStmt><distStmt><distrbtr source="archive">DataverseNO</distrbtr><distDate>2024-11-26</distDate></distStmt><verStmt source="archive"><version date="2025-07-17" type="RELEASED">1</version></verStmt><biblCit>Sönning, Lukas, 2024, "Background data for: Advancing our understanding of dispersion measures in corpus research", https://doi.org/10.18710/FVHTFM, DataverseNO, V1</biblCit></citation></docDscr><stdyDscr><citation><titlStmt><titl>Background data for: Advancing our understanding of dispersion measures in corpus research</titl><IDNo agency="DOI">doi:10.18710/FVHTFM</IDNo></titlStmt><rspStmt><AuthEnty affiliation="University of Bamberg">Sönning, Lukas</AuthEnty></rspStmt><prodStmt><producer>University of Bamberg</producer><prodDate>2023-06-28</prodDate><prodPlac>Bamberg, Germany</prodPlac><software version="22.5.0">MAXQDA Plus</software><software version="4.2.1">R</software></prodStmt><distStmt><distrbtr source="archive">DataverseNO</distrbtr><distrbtr abbr="TROLLing" URI="https://trolling.uit.no/">The Tromsø Repository of Language and Linguistics (TROLLing)</distrbtr><contact affiliation="University of Bamberg" email="lukas.soenning@uni-bamberg.de">Sönning, Lukas</contact><depositr>Sönning, Lukas</depositr><depDate>2023-12-19</depDate></distStmt><holdings URI="https://doi.org/10.18710/FVHTFM"/></citation><stdyInfo><subject><keyword xml:lang="en">Arts and Humanities</keyword><keyword>dispersion</keyword><keyword>corpus linguistics</keyword><keyword>methodology</keyword><keyword>corpus design</keyword><keyword>Brown Corpus</keyword><keyword>dispersion measures</keyword><keyword>lexical dispersion</keyword><keyword>word importance</keyword><keyword>vocabulary lists</keyword><keyword>word frequency lists</keyword><keyword>text-level analysis</keyword><keyword>frequency</keyword><keyword>Juilland's D</keyword><keyword>Gries' DP</keyword><keyword>DA</keyword><keyword>English</keyword></subject><abstract date="2023-12-19">&lt;p>&lt;b>Dataset description&lt;/b>&lt;/p>
&lt;p>This dataset contains background data and supplementary material for Sönning (forthcoming), a study that looks at the behavior of dispersion measures when applied to text-level frequency data. For the literature survey reported in that study, which examines how dispersion measures are used in corpus-based work, it includes tabular files listing the 730 research articles that were examined as well as annotations for those studies that measured dispersion in the corpus-linguistic (and lexicographic) sense. As for the corpus data that were used to train the statistical model parameters underlying the simulation study reported in that paper, the dataset contains a term-document matrix for the 49,604 unique word forms (after conversion to lower-case) that occur in the Brown Corpus. Further, R scripts are included that document in detail how the Brown Corpus XML files, which are available from the Natural Language Toolkit (Bird et al. 2009; https://www.nltk.org/), were processed to produce this data arrangement.&lt;/p></abstract><abstract date="2023-12-19">&lt;p>&lt;b>Abstract: Related publication&lt;/b>&lt;/p>
&lt;p>This paper offers a survey of recent corpus-based work, which shows that dispersion is typically measured across the text files in a corpus. Systematic insights into the behavior of measures in such distributional settings are currently lacking, however. After a thorough discussion of six prominent indices, we investigate their behavior on relevant frequency distributions, which are designed to mimic actual corpus data. Our evaluation considers different distributional settings, i.e. various combinations of frequency and dispersion values. The primary focus is on the response of measures to relatively high and low sub-frequencies, i.e. texts in which the item or structure of interest is over- or underrepresented (if not absent). We develop a simple method for constructing sensitivity profiles, which allow us to draw instructive comparisons among measures. We observe that these profiles vary considerably across distributional settings. While D and DP appear to show the most balanced response contours, our findings suggest that much work remains to be done to understand the performance of measures on items with normalized frequencies below 100 per million words.&lt;/p></abstract><sumDscr><timePrd cycle="P1" event="start" date="1961-01-01">1961-01-01</timePrd><timePrd cycle="P1" event="end" date="1961-12-31">1961-12-31</timePrd><collDate cycle="P1" event="start" date="2023-06-14">2023-06-14</collDate><collDate cycle="P1" event="end" date="2023-06-28">2023-06-28</collDate><nation>United States</nation><dataKind>textual linguistic data</dataKind><dataKind>corpus data</dataKind><dataKind>observational data</dataKind></sumDscr></stdyInfo><method><dataColl><sources><dataSrc>&lt;p>A Standard Corpus of Present-Day Edited American English, for use with Digital Computers (the Brown Corpus). 1964, 1971, 1979. Compiled by W. N. Francis and H. Kučera. Brown University. Providence, Rhode Island.&lt;/p> 
&lt;p>Brown Corpus XML files are available from the Natural Language Toolkit (&lt;a href="https://www.nltk.org">https://www.nltk.org&lt;/a>).&lt;/p>
&lt;p>The extracted words included in the data files of this dataset represent insubstantial portions of the Brown Corpus; they do not represent coherent stretches of text. Reuse of such excerpts is permitted under exceptions in IPR and database protection regulations, such as the Norwegian Copyright Act (cf. &lt;a href="https://lovdata.no/lov/2018-06-15-40/§24">§ 24 Eneretten til databaser&lt;/a>), the &lt;a href="http://data.europa.eu/eli/dir/1996/9/oj">EU Database Directive&lt;/a> (cf. art 8 Rights and obligations of lawful users), and Fair use (cf. &lt;a href="https://www.copyright.gov/fair-use/more-info.html">US Copyright Act&lt;/a>).&lt;/p></dataSrc></sources></dataColl><anlyInfo/></method><dataAccs><setAvail/><useStmt/><notes type="DVN:TOU" level="dv">&lt;p>With the exception of the tab-delimited term-document matrix brown_tdm.tsv and the tab-delimited document-term matrix brown_dtm.tsv, the dataset "Background data for: Advancing our understanding of dispersion measures in corpus research" has been marked as dedicated to the public domain, as described here: &lt;a href="https://creativecommons.org/publicdomain/zero/1.0/">https://creativecommons.org/publicdomain/zero/1.0/&lt;/a>.&lt;/p>
&lt;p>Our Community Norms as well as good scientific practices expect that proper credit is given via citation. Please use the data citation shown on the dataset page.&lt;/p>
&lt;p>The tab-delimited term-document matrix brown_tdm.tsv and the tab-delimited document-term matrix brown_dtm.tsv contain word forms that have been extracted from the Brown Corpus, available from the Natural Language Toolkit (&lt;a href="https://www.nltk.org/">https://www.nltk.org/&lt;/a>), under limitations and exceptions to IPR and database protection regulations. The contribution of the author of the present dataset to these files, as detailed in the ReadMe file, is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, as described here: 
&lt;a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/&lt;/a>. Reusers should note that this license does not apply to the word forms extracted from the Brown Corpus.&lt;/p>
</notes></dataAccs><othrStdyMat><relPubl><citation><titlStmt><titl>Sönning, Lukas. 2025. Advancing our understanding of dispersion measures in corpus research. Corpora 20(1). 3-35.</titl><IDNo agency="doi">10.3366/cor.2025.0326</IDNo></titlStmt><biblCit>Sönning, Lukas. 2025. Advancing our understanding of dispersion measures in corpus research. Corpora 20(1). 3-35.</biblCit></citation><ExtLink URI="https://doi.org/10.3366/cor.2025.0326"/></relPubl></othrStdyMat></stdyDscr><otherMat ID="f233085" URI="https://dataverse.no/api/access/datafile/233085" level="datafile"><labl>00ReadMe_understanding_dispersion.txt</labl><txt>File describing the dataset</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">text/plain</notes></otherMat><otherMat ID="f192970" URI="https://dataverse.no/api/access/datafile/192970" level="datafile"><labl>2023-12-08_dispersion_survey_all_articles.tsv</labl><txt>Tab-delimited data table containing the 730 research articles that entered our literature survey</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">text/tsv</notes></otherMat><otherMat ID="f192971" URI="https://dataverse.no/api/access/datafile/192971" level="datafile"><labl>2023-12-09_dispersion_survey.tsv</labl><txt>Tab-delimited data table containing annotations for the 38 studies in our survey that assessed dispersion</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">text/tsv</notes></otherMat><otherMat ID="f233028" URI="https://dataverse.no/api/access/datafile/233028" level="datafile"><labl>brown_dtm.tsv</labl><txt>Tab-delimited document-term matrix for the word forms in the Brown Corpus</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">text/tsv</notes></otherMat><otherMat ID="f233029" URI="https://dataverse.no/api/access/datafile/233029" level="datafile"><labl>brown_tdm.tsv</labl><txt>Tab-delimited term-document matrix for the word forms in the Brown Corpus</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">text/tsv</notes></otherMat><otherMat ID="f233030" URI="https://dataverse.no/api/access/datafile/233030" level="datafile"><labl>script_brown_data_retrieval.qmd</labl><txt>R quarto script documenting retrieval of the data from the Brown XML files</txt><notes level="file" type="DATAVERSE:CONTENTTYPE" subject="Content/MIME Type">application/octet-stream</notes></otherMat></codeBook>