Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Performance and reliability" (doi:10.18710/EUXSMW)

View:

Part 1: Document Description
Part 2: Study Description
Part 5: Other Study-Related Materials
Entire Codebook

(external link)

Document Description

Citation

Title:

Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Performance and reliability"

Identification Number:

doi:10.18710/EUXSMW

Distributor:

DataverseNO

Date of Distribution:

2023-04-18

Version:

1

Bibliographic Citation:

Sönning, Lukas, 2023, "Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Performance and reliability"", https://doi.org/10.18710/EUXSMW, DataverseNO, V1

Study Description

Citation

Title:

Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Performance and reliability"

Identification Number:

doi:10.18710/EUXSMW

Authoring Entity:

Sönning, Lukas (University of Bamberg)

Producer:

University of Bamberg

Distributor:

DataverseNO

Distributor:

The Tromsø Repository of Language and Linguistics (TROLLing)

Access Authority:

Sönning, Lukas

Depositor:

Sönning, Lukas

Date of Deposit:

2022-07-08

Holdings Information:

https://doi.org/10.18710/EUXSMW

Study Scope

Keywords:

Arts and Humanities, keyness, keywords, corpus, methodology, frequency, corpus linguistics, dispersion, COCA, English, Corpus of Contemporary American English

Abstract:

This dataset contains corpus-based frequency data for an analysis of key verbs in published academic writing. The data are from the Corpus of Contemporary American English (COCA; Davies 2008-) and cover a period of 30 years (1990-2019). The section ‘academic’, which contains research articles from peer-reviewed journals, represents the target variety, and the reference variety is fictional writing as represented in the ‘fiction’ section (which contains short stories, plays, movie scripts, and the first chapter of novels). The total number of text files is 26,137 (academic) and 25,992 (fiction). To reduce computational expense for our methodological simulation study, we restrict our attention to verb lemmas whose whole-(sub)corpus normalized frequency exceeds 10 pmw in the academic section of COCA. The data therefore contain frequency information on only 700 verb lemmas.

Time Period:

1990-01-01-2019-12-31

Date of Collection:

2008-2020

Country:

United States

Kind of Data:

corpus data

Kind of Data:

observational data

Methodology and Processing

Sources Statement

Data Sources:

COCA (Corpus of Contemporary American English); Davies, Mark. 2008. The Corpus of Contemporary American English. www.english-corpora.org/coca.

Data Access

Notes:

<p>This Dataset, Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Reliability and interpretability" (hereafter: Dataset), may be reused according to the <strong>CLARIN PUB+BY+LRT</strong> license as described below:</p> <p></p> <p><strong>Copyright holder: </strong>Lukas Sönning</p> <p></p> <p><strong>Resource:</strong> Key verbs in academic writing: Dataset for "Evaluation of keyness metrics: Reliability and interpretability"</p> <p></p> <p>The Copyright holders grant the End-User a free, non-exclusive and perpetual (for the duration of the copyright) right to use and make copies of the Resource, distribute copies and present the Resource in public as such, as modified, or as part of a compilation or derived work. The permission applies to all known or future modes and means of communication and includes a right to make modifications enabling the use of the Resource on other devices and in other formats.</p> <p></p> <p>Additional license terms as defined in the <em>Condition Definitions</em> section below:<br> <ul> <li><strong>General Use conditions:</strong> BY, LRT</li> </ul> <p></p> <p><strong>Condition Definitions:</strong><br> <ul> <li><strong>BY:</strong> Attribution, i.e. acknowledgement of authorship, is required.</li> <li><strong>LRT:</strong> The content is available only for language research and technology development.</li> </ul></p> <p></p> <p>This license has been made in compliance with copyright agreements by WIPO – the World Intellectual Property Organization. The rights granted in this license shall be so interpreted that in case applicable intellectual property laws grant rights not mentioned in this license, they are also regarded as part of the rights to be licensed; the purpose of this license is not to restrict any rights intended to be licensed within different legal systems. Additional rights to the Resource may be agreed separately in writing.</p>

Other Study Description Materials

Related Materials

Sönning, Lukas. 2023. Evaluation of keyness metrics: Performance and reliability. Open Science Framework project. https://doi.org/10.17605/OSF.IO/KCWUS

Related Publications

Citation

Title:

Sönning, Lukas. 2024. Evaluation of keyness metrics: Performance and reliability. Corpus Linguistics and Linguistic Theory 20(2). 263–288.

Identification Number:

10.1515/cllt-2022-0116

Bibliographic Citation:

Sönning, Lukas. 2024. Evaluation of keyness metrics: Performance and reliability. Corpus Linguistics and Linguistic Theory 20(2). 263–288.

Other Study-Related Materials

Label:

00_ReadMe_keyverbs.txt

Text:

Documentation file describing the dataset

Notes:

text/plain

Other Study-Related Materials

Label:

2022-12-21_text_metadata_for_coca2020.txt

Text:

Metadata for the text files in COCA 2020

Notes:

text/plain

Other Study-Related Materials

Label:

2022-12-22_coca_metadata_ACAD_FIC.txt

Text:

Metadata for the text files in the dataset

Notes:

text/plain

Other Study-Related Materials

Label:

2022-12-22_verb_lemmas_acad.txt

Text:

Text-level occurrences of the 700 verb lemmas (COCA section "academic")

Notes:

text/plain

Other Study-Related Materials

Label:

2022-12-22_verb_lemmas_fict.txt

Text:

Text-level occurrences of the 700 verb lemmas (COCA section "fiction")

Notes:

text/plain

Other Study-Related Materials

Label:

script_data_preparation.qmd

Text:

R script documenting the preparation of the data for analysis

Notes:

application/octet-stream

Other Study-Related Materials

Label:

script_data_retrieval.qmd

Text:

R script documenting data retrieval from the full-text corpus files (COCA)

Notes:

application/octet-stream